=== AI Crawler Log and Content Signals by TheSEO ===
Contributors: theseo
Tags: ai, crawler, robots txt, logging, seo
Requires at least: 6.0
Tested up to: 7.0
Requires PHP: 7.4
Stable tag: 1.0.0
License: GPLv2 or later
License URI: https://www.gnu.org/licenses/gpl-2.0.html

See which AI crawlers fetch your pages, and state in robots.txt what they may do with the content. Self hosted, no outgoing request.

== Description ==

Two sides of the same question, in one plugin.

**Which AI crawlers came by.** A log of every fetch by a crawler from a bundled
list of 34 known agents, with the page, the moment and the HTTP status your
server returned. Counted per crawler and per page, over 7, 30 or 90 days
or over everything, with a CSV export.

**What they may do with it.** A screen where you state, per crawler group,
whether you allow your content to be used for search, as AI input, and for AI
training. Those three statements are written into robots.txt as content
signals. You see the exact lines before you save, and nothing is added until
you switch it on.

Everything stays in your own WordPress. The plugin makes no outgoing request at
all, has no account, no dashboard elsewhere and no external service behind it.

= What it stores =

One row per fetch:

* the name of the crawler, taken from the list in the plugin
* the address of the page, without the query string and without the fragment
* the HTTP status your server returned
* the server timestamp

= What it does not store =

* no IP address, not in plain text and not as a hash
* no user agent string, only the name it matched
* no cookie and nothing in localStorage or sessionStorage
* no user id
* no query string, so a token or an e-mail address in a url never reaches the
  table
* nothing that leaves your server

The visitor address is used once, to count requests per minute so a spoofed
user agent cannot flood the table. It becomes a keyed hash in a value that
expires within a minute and it never reaches the log table.

= What this plugin does not claim =

It does not say whether your site is quoted in ChatGPT, Perplexity, Gemini or
any other assistant. That cannot be measured from your own server and the
plugin therefore does not pretend to.

A crawler is recognised by the user agent it sends, and any client can claim
any user agent. The log tells you what was claimed, not who it really was.
There is no reverse DNS or IP range verification in this version.

Two entries in the list, Google-Extended and Applebot-Extended, exist only as
robots.txt groups. They are never sent in a request, so you can give them a
signal but they can never appear in the log. The plugin says so on screen
rather than quietly counting zero.

= Content signals =

A content signal states what you prefer a crawler to do with your content
after it has fetched it. Three signals exist:

* `search` for building a search index and showing links and short excerpts
* `ai-input` for feeding content into a model at answer time
* `ai-train` for training or fine tuning a model

Each one can be set to yes, to no, or to no statement. No statement leaves the
signal out of the line entirely, which is how the policy expresses that you
neither grant nor restrict that use. Everything starts at no statement and the
robots.txt addition starts switched off, so installing this plugin never
changes what crawlers are told until you decide it should.

A signal is a statement of preference, not a technical block. It depends on the
other party honouring it. Next to the three signals there is a separate column
that adds a real `Disallow: /` for a group, which is a different thing and is
off everywhere until you switch it on.

The standardisation of these signals is still being worked out in the AIPREF
working group of the IETF. The plugin writes them in the form published by the
Content Signals Policy, https://contentsignals.org/, and will follow that
wording as it settles.

= Retention =

A daily job deletes rows older than the retention setting, 90 days by default.
Deactivating the plugin keeps the log and stops the job. Deleting the plugin
drops the table, the settings and the job.

== Installation ==

1. Upload the plugin folder to `/wp-content/plugins/theseo-ai-crawlerlog`, or
   upload the zip through Plugins, Add new, Upload plugin.
2. Activate the plugin. Logging starts straight away, the robots.txt side
   stays off.
3. Open AI Crawlers, Settings and check the retention and the crawler list.
4. Open AI Crawlers, Content signals when you want to state what crawlers may
   do with your content. Set the signals, look at the preview, then tick the
   box that adds the block to robots.txt and save.
5. Come back to AI Crawlers, Log after a few days. Crawlers do not visit on
   command.

== Frequently Asked Questions ==

= Does anything leave my server? =

No. The plugin makes no outgoing HTTP request. Everything it knows comes from
requests that arrived at your own site, and everything it stores goes into your
own database table.

= My log stays empty, is it broken? =

Maybe not. Three ordinary reasons. Crawlers may simply not have come by yet on
a small or new site. Logging may be switched off on the settings screen. Or the
requests never reach PHP because a full page cache, a reverse proxy or a CDN
answers them, in which case no WordPress plugin can see them. The settings
screen prints the list of user agents so you can paste it into the exclusion
list of your caching plugin.

= Does it slow the site down? =

The check runs on requests that have a user agent, compares it against a list
of strings, and stops there for a normal visitor. Only a matching crawler
causes one INSERT. Nothing is added to the page, no script is loaded and no
visitor notices anything.

= Why is there no nonce anywhere in the logging? =

Because the logging never accepts input. It looks at requests that arrive by
themselves and writes a row on the server side. There is no endpoint to post
to. The admin screens, the export and every save do use a nonce and a
capability check.

= Does the robots.txt block always show up? =

Only when WordPress generates robots.txt. If a real robots.txt file sits in the
root of your site, your web server serves that file and WordPress never gets a
say. The signals screen checks for that file and tells you when it is there.

= Can I add a crawler that is not in the list? =

Yes. The settings screen has a box for your own tokens, one per line. Each line
is looked for in the user agent header, so a token like `MyBot` matches
`MyBot/1.0 (+https://example.com/bot)`.

= Does it work on multisite? =

Yes. Each site in the network gets its own table, its own settings and its own
robots.txt block. Deleting the plugin cleans up every site.

= Can I use it on client sites? =

Yes. The data stays in the client's own WordPress, the settings screen states
exactly what is stored, and there is nothing to show a data protection officer
beyond that screen, because no personal data is kept.

== Changelog ==

= 1.0.0 =
* First version.
* Log of fetches by 34 known AI crawlers, 32 of which can appear in a request,
  per crawler and per page, over 7, 30 or 90 days or over everything.
* Four totals, a daily chart and a CSV export of at most 20000 lines.
* Content signals for `search`, `ai-input` and `ai-train` per crawler group,
  three states each, written into robots.txt through the WordPress filter, with
  a preview of the exact lines before saving.
* An optional `Disallow: /` per group, off by default and clearly separated
  from the signals.
* Warnings when a real robots.txt file exists or when the site discourages
  search engines, because in both cases the block does not do what it looks
  like it does.
* Retention of 30, 90 or 365 days with a daily cleanup job, or keep everything.
* Your own user agent tokens, at most 25.
* No IP address, no user agent string, no cookie, no browser storage and no
  outgoing request.

== Upgrade Notice ==

= 1.0.0 =
First version.
