AI Content Signals
AI Content Signals lets you declare, in a machine-readable way, how AI systems may use your content: for search indexing, real-time AI answers (RAG), or model training. It started with Cloudflare’s Content Signals in robots.txt and now expresses the same preferences across several surfaces, so more crawlers and tools can read them: robots.txt — Cloudflare Content Signals (search / ai-input / ai-train) HTTP header — X-Robots-Tag: noai, noimageai HTML meta — robots noai, noimageai /.well-known/tdmrep.json — W3C Text and Data Mining Reservation Protocol EU Directive 2019/790 rights reservation Everything is opt-in and does not change your robots.txt output unless you enable it. The signals you control You set three preferences, and AI Content Signals expresses each one in the right format on every surface you enable: search — allow or deny search indexing and traditional search results (links and short snippets) ai-input — allow or deny using your content for real-time AI answers (RAG, grounding, AI Overviews) ai-train — allow or deny using your content to train or fine-tune AI models These three come from Cloudflare’s Content Signals vocabulary, written to your robots.txt. When you deny AI training, the same opt-out is also emitted on any other surface you enable: noai, noimageai in an HTML meta tag and an X-Robots-Tag header, plus a tdm-reservation in your TDMRep manifest. Declaring the same preference in several places means a crawler that ignores one signal may still honor another. Key Features Easy-to-use settings page in WordPress admin Set global defaults for all crawlers Configure specific settings for individual AI bots (GPTBot, ClaudeBot, PerplexityBot, etc.) Add custom bot User-Agents Supports both physical and virtual robots.txt files Option to create physical robots.txt with basic WordPress rules Preview generated Content Signals before applying Export and import settings as JSON for easy migration between sites Optional legal text with EU Directive reference Developer-friendly: filter hook to extend the predefined bots list Works with existing robots.txt from SEO plugins Automatic sitemap detection and inclusion Optional extra output surfaces: HTML robots meta tag (noai, noimageai), X-Robots-Tag header, and a W3C TDMRep file at /.well-known/tdmrep.json Supported Bots The plugin includes predefined settings for 28 major AI crawlers: OpenAI GPTBot, OAI-SearchBot, and ChatGPT-User Anthropic ClaudeBot, Claude-Web, and anthropic-ai Perplexity Bot and Perplexity-User Google Extended (Gemini) and GoogleOther Amazon Bot Apple Extended Meta/Facebook Bot and meta-externalagent DuckDuckGo DuckAssistBot Allen Institute AI2Bot Mistral AI ByteDance Bytespider DeepSeek AI xAI Grok Huawei Pangu Common Crawl, Cohere AI, Diffbot, You.com Bot, and more Important Notice Content Signals is a declarative standard – it expresses your preferences but does not technically enforce them. AI companies are not legally required to respect these signals, though the plugin includes legal text referencing EU copyright directives. The IETF AI Preferences (AIPREF) Working Group is currently developing a formal standard based on similar concepts. This plugin implements the current Cloudflare Content Signals specification and will be updated as standards evolve. This plugin works best when combined with other protection measures like traditional robots.txt rules and server-level bot management. Support Need private support or custom development? Do you need one-on-one help, priority troubleshooting, or a custom feature, integration, or tweak built specifically for your site? I offer private support and custom development. Just contact me and tell me what you need. Need help or have suggestions? Official website WordPress support forum YouTube channel Documentation and tutorials Love the plugin? Please leave us a 5-star review and help spread the word! About AyudaWP.com We are specialists in WordPress security, SEO, and performance optimization plugins. We create tools that solve real problems for WordPress site owners while maintaining the highest coding standards and accessibility requirements.
Top keywords
- ai15×2.46%
- content11×1.80%
- robots10×1.64%
- signals10×1.64%
- content signals8×1.31%
- robots txt8×1.31%
- txt8×1.31%
- bot6×0.98%
- search5×0.82%
- wordpress5×0.82%
- cloudflare4×0.66%
- custom4×0.66%
AI Scrape Protect
AI Scrape Protect is a WordPress plugin designed to protect your website from scraping for AI training purposes. It achieves this by adding opt-out instructions to the robots.txt file for the most common AI scraping bots and including meta tags to control how your content is used. Version 5.0 introduces a full settings page with granular per-bot control, categorised bot management, AJAX-powered toggles, and support for custom user agents. Features Adds User-agent and Disallow rules to your robots.txt file to block a comprehensive list of AI scraping bots. Generates compact robots.txt output: all blocked bots are grouped under a single Disallow: / directive. Introduces meta tags in the HTML to provide additional instructions to AI bots. Separate toggles for content meta tags (noai, nosummary, DisallowAITraining) and image meta tags (noimageai). Settings page with five categorised tabs: Search Engines, AI Training, AI Search & Answers, General Indexing, Custom Bots. Group toggle per category to block or allow all bots in that category at once. Individual per-bot toggle with description and company information for each bot. Custom bots tab: add and remove your own user agent strings. Global enable/disable toggle and separate toggles for robots.txt and meta tags. “Apply default settings” button to reset all options. Admin bar icon linking to the settings page (visible to administrators only). Icon reflects the current enabled state. Physical robots.txt detection notice: warns you if a file in your root directory overrides WordPress’s virtual robots.txt. Fully AJAX-powered settings page: changes are saved instantly without page reloads. Extensible via the aisp_user_agents filter for developers. Multisite note: settings are stored per site using get_option/update_option. Network-wide configuration is not supported. Note: robots.txt and meta tag opt-out instructions are not always respected by all bots. This plugin is a measure to discourage scraping rather than a guaranteed technical block.
Top keywords
- bots8×2.56%
- ai7×2.24%
- robots7×2.24%
- robots txt7×2.24%