AI crawler log and content signals for WordPress
GPTBot fetches your page, so does ClaudeBot, and your server log only tells you once you go digging yourself. This plugin puts those two questions on one screen: which AI crawler has been by and on which page, and what they're allowed to do with what they fetch. That second part you set in your robots.txt, with the exact rules shown on screen before you save. The plugin makes no outgoing requests of its own and logs no IP address.
Three steps, and for an ordinary visitor it stops at step two
The plugin hooks into the requests that arrive at your server anyway. Nothing is added to your page, no script is loaded, and a visitor notices nothing. This is what happens, and nothing more.
A request comes in
WordPress receives a request for a page. The plugin looks at one thing: the user agent the other side sends along.
If the request has no user agent at all, it stops here.
Compared with the list
The user agent is compared with a list of 34 names built into the plugin itself, plus any names you've typed in yourself.
No match, and nothing further happens. That's the path for every ordinary visitor.
A row added
A match, and exactly one INSERT follows with four fields: the crawler's name, the page's address, the status code your server returned and the time.
The table has no other fields.
Your server log knows, you just don't read it
Every AI crawler that fetches your site leaves a trace in your web server's access log. In practice nobody looks. The file is huge, full of images and bot traffic, and most hosting packages keep it for only a few days. The question you actually have is small: did GPTBot come by, on which pages, and did my server return a 200 or a 404.
Since 2025 there's a second question that didn't used to exist
A crawler fetching your page can do three different things with it: put it in a search index, use it as a source while answering a question, or use it to train a model. Those are three separate forms of use, and until recently you had no way to say something different per form. Disallow was all or nothing.
This plugin does those two things side by side, because in practice it's a conversation
You see who's been by, and on that basis you decide what you allow.
What the plugin actually does
Two screens, and a third with the settings.
The log
One row every time a crawler from the list fetched a page on your site. Counted per crawler and per page, over 7, 30 or 90 days or over everything, with four totals at the top, a bar per day and an export to CSV of up to 20,000 rows.
The content signals
A screen where you state, per crawler group, whether you make your content available for search, as input for an AI answer, and for training. You see the exact rules that follow, spelled out on screen, before you save, and nothing is added to your robots.txt until you tick the box yourself. There's also a separate column that adds a genuine Disallow: / for a group. That's a different thing from a signal, it's off everywhere by default, and the screen keeps the two apart.
The settings
Measuring on or off, a retention period of 30, 90 or 365 days or keep everything, and a field for your own names, up to 25, one per line. Each line is searched for within the user agent, so MyBot also matches MyBot/1.0 (+https://example.com/bot).
Three statements per group, and you see the rules before you save
A content signal is a statement about what you allow with your content once it's been fetched. There are three of them and they're independent of each other. Each signal can be yes, no, or no stated preference. No stated preference leaves the signal out of the line entirely, and that's exactly how you say you neither allow nor forbid that form of use. Everything starts on no stated preference, and the block is off.
robots.txt file at the root of your site, your web server serves that file and WordPress never gets involved. The screen checks for this and tells you.How to install it, in four steps
No account, no key and no registration needed. An ordinary WordPress site where you're an administrator is enough.
manage_options, so administrator.Click the download button above. You get theseo-ai-crawlerlog-1.0.0.zip, 52 KB. Don't unzip it: WordPress wants the zip itself.
In your WordPress: Plugins, Add New Plugin, Upload Plugin. Choose the file and click Install Now.
Click Activate Plugin. A new item appears in your admin menu. Your site doesn't change yet at this point.
Open the settings screen and go through the options. Anything that touches your public site is off until you switch it on yourself.
What it doesn't do, and doesn't claim
This is here because it's the fastest way to get this topic wrong.
It doesn't tell you whether you're quoted
Whether ChatGPT, Perplexity or Gemini mentions your site in an answer can't be measured from your own server. A crawler visit isn't a citation, and the plugin doesn't pretend it is.It doesn't verify the crawler
A user agent is a claim. Anyone posing as GPTBot ends up in the log as GPTBot. This version has no reverse DNS check and no IP range check, and the screen says so too.Two names can never appear in the log
Google-Extended and Applebot-Extended only exist as a group in robots.txt and are never sent with a request. You can give them a signal, but they'll never show up in the list. The plugin says so on screen rather than silently counting zero.An empty log usually isn't a fault
Three common causes: no crawler has been by yet, measuring is switched off, or a full-page cache or CDN answers the requests so they never reach PHP. In that last case almost no WordPress plugin can see it, unless it reads your CDN's log afterwards. The settings screen prints out the list of names so you can paste it into your cache plugin's exceptions.Four fields in, and six things deliberately left out
This is the whole table. There's no second table, no hidden column and no second place anything ends up. The right-hand column isn't a promise but the reason there's no personal data in this log that you'd need to write into a record of processing activities.
- The crawler's name, as it appears in the plugin's list
- The page's address, without a query string and without an anchor
- The HTTP status code your server returned
- The time, according to your server
- The IP address, not in readable form and not as a fingerprint either
- The full user agent string
- A cookie or anything in localStorage or sessionStorage
- A user number or a link to a WordPress account
- The query string, so a key or email address in an address never ends up in it
- Anything at all that leaves your server
includes/class-theseo-crawlerlog-request.php.Privacy: what is and isn't recorded
What goes into the table is set out above in FIG.03, field by field. What follows here is how you check it yourself, because a privacy paragraph you can't verify is just marketing copy.
No outgoing requests at all
Search the unzipped file forwp_remote_, curl_init, file_get_contents and fsockopen. Zero matches. There's no account, no key and no service behind it.No IP address in the table
The address appears exactly once, in the rate limit, where it becomes a transient name via a key fingerprint that expires within a minute.No cookie and no browser storage
The plugin loads no script at all on your front end. There's nothing that sets or reads anything on a visitor's device.No query string
The address is rebuilt from scratch, so a session token or an email address someone put in an address never ends up in the table.Retention period
A daily job deletes rows older than the set period, 90 days by default. Deactivating keeps the log and stops the job. Deleting drops the table, the settings and the job, readable inuninstall.php.For the sake of completeness: this log is about machines, not people. It contains no data traceable to a person, which is why you can put it on a client site without needing anything added to a record of processing activities.
Frequently asked questions
Does the plugin slow down my site? The check runs on requests that carry a user agent, compares it with a list of short strings, and stops there for an ordinary visitor. Only a match produces an INSERT. Nothing is added to the page and no script is loaded.
Why isn't there a nonce on the logging?
Because logging doesn't accept input. It looks at requests that arrive on their own and writes a row server side. There's no address to send anything to. The admin screens, the export and every save action do use a nonce plus a permissions check.
Does it work with a cache plugin?
Only to the extent that requests reach PHP. If your cache or your CDN answers the request itself, WordPress never sees it, and so neither does this plugin. The settings screen prints out the list of names so you can exclude them from the cache.
Does it work on multisite?
Yes. Every site in the network gets its own table, its own settings and its own robots.txt block. Deleting cleans up every site.
Can I put it on client sites?
Yes. The data stays within that client's WordPress, and the settings screen states plainly what's kept. There's no personal data to show a data protection officer.
Does it change my robots.txt as soon as I activate it?
No. Logging starts straight away; the robots.txt side is off. Nothing is added until you set the signals, review the preview and tick the box.
What you get out of the first week
Turn it on for a week and then see who's been by. That conversation is almost always shorter and more concrete than the conversation about AI visibility in general. Want to know afterwards how to make those pages citable too, not just fetchable? The AI snippet previewer is the next step, and the llms.txt generator states in plain language what your site is. Want to talk it through, book a call.