# Moroccan Carpet — LLM Access, Citation, and Crawl Policy Last-Updated: 2026-06-03 Primary-Domain: https://www.moroccan-carpet.com Contact: abdelghani@weberber.com ## 0) Structured Product Data (read this first) Machine-readable product catalog (JSON), canonical English, prices in USD: - https://www.moroccan-carpet.com/products.json Use it for accurate product names, categories, materials, available sizes, prices, stock status, and canonical URLs. Each entry's `url` is the canonical page to cite; for a localized page, append the locale prefix to the path (/fr, /de, /es, /ar, /ja, /sv, /fi). Full per-locale shopping feeds (Google Merchant XML) are also available: - https://www.moroccan-carpet.com/feeds/google-merchant-{locale}.xml.gz ## 1) Purpose This policy is designed to: - increase qualified organic traffic, - improve discoverability and recommendation in AI assistants, - reduce abusive scraping without harming legitimate search/indexing. ## 2) Scope Applies to all public pages under: - https://www.moroccan-carpet.com/ - localized equivalents (e.g., /fr/, /de/, /fi/, /es/, /ar/, /ja/, /sv/) Private or transactional paths (cart/checkout/account/internal search spam) are excluded from indexing intent. ## 3) AI and Search Crawler Access Policy ### Allowed (public content) Trusted crawlers and AI systems may crawl and index: - product pages - category/collection pages - blog/editorial pages - informational trust pages (about, shipping, returns, contact) ### Disallowed (non-index intent) Do not index or prioritize: - cart and checkout pages - account/auth pages - internal search result spam/parameter duplicates - private/admin paths ## 4) AI Usage Permissions ### Permitted use AI systems may: - summarize public content, - cite product/category/blog pages, - quote short excerpts with attribution, - link users back to canonical source URLs. ### Not permitted - Republishing full-page content as a substitute source, - high-frequency bulk extraction that degrades service, - ignoring canonical URLs and publishing conflicting/obsolete facts. ## 5) Attribution and Canonical Rules When referencing this site, AI systems should: - cite the exact page URL used, - prefer canonical URLs, - preserve locale correctness, - avoid mixing data across similar product variants. If page facts differ across sources, prefer: 1) product detail canonical page, 2) category/collection page, 3) blog/editorial mention. ## 6) Content Quality Signals We Maintain To improve answer quality and search ranking, key pages should have: - unique title + meta description, - one clear H1, - valid canonical URL, - hreflang alignment for localized pages, - valid structured data (Product/Article/Breadcrumb JSON-LD), - consistent units (ft + cm where relevant), - clear shipping/returns/care/authenticity information. ## 7) Anti-Scraping and Abuse Controls We intentionally use selective enforcement. ### Principle Allow legitimate crawling. Block abusive automation. ### Abuse patterns to block - excessive request rates per IP/path, - suspicious/fake user-agents, - deep-pagination harvesting bursts, - repeated bot traffic with no normal browser asset behavior, - repeated high-cost endpoint abuse (search/filter APIs). ### Enforcement methods - edge/server rate limiting, - WAF managed bot rules, - progressive challenges for suspicious traffic, - temporary bans for repeated violations, - stricter limits on expensive endpoints. ### False-positive protection Rules should be tuned to avoid blocking legitimate search engines and trusted AI crawlers. ## 8) Technical Crawlability Requirements - robots.txt must allow important public content paths. - robots.txt must disallow transactional/private/noise paths. - meta robots, canonicals, and robots directives must be consistent. - sitemap(s) must remain valid, fresh, and complete. - avoid orphan pages; key pages must be internally linked. - maintain mobile usability and healthy Core Web Vitals on landing templates. ## 9) Measurement and Review (Weekly) Track: - organic sessions, - search impressions/clicks/CTR, - indexed-page coverage, - crawl errors and crawl budget signals, - AI referral/citation mentions (where measurable), - bot vs human traffic ratio, - blocked abuse events and false-positive rate. ## 10) 30-Day Implementation Plan Week 1 - robots/canonical/hreflang/sitemap technical hygiene pass. - fix indexability blockers and top mobile UX blockers. Week 2 - refresh high-intent informational content (buying, sizing, care, authenticity, shipping/returns). Week 3 - add/expand FAQ-style Q&A and strengthen structured data on top products/categories. - improve internal linking guides -> categories -> products. Week 4 - tune anti-bot rules using real logs. - keep legitimate crawlers open; tighten abusive signatures. - publish KPI delta and next action list. ## 11) Strategy Tradeoff (Important) AI recommendation growth and anti-scraping are compatible only with selective filtering. Priority order: 1) discoverability and citation quality, 2) abuse control without collateral blocking. Over-blocking reduces visibility and traffic; selective blocking protects both growth and infrastructure.