EN
Language · same page NLNederlands/tools-en-plugins/ai-crawler-log/ ENEnglish (UK)/en/tools-and-plugins/ai-crawler-log/ ESEspañolnot translated yet We do not remember your choice and never redirect you automatically.
SRV · plugin · ai crawlers reg. T.000 · wordpress · gplv2 or later

AI crawler log and content signals for WordPress

GPTBot fetches your page, so does ClaudeBot, and your server log only tells you once you go digging yourself. This plugin puts those two questions on one screen: which AI crawler has been by and on which page, and what they're allowed to do with what they fetch. That second part you set in your robots.txt, with the exact rules shown on screen before you save. The plugin makes no outgoing requests of its own and logs no IP address.

Download the plugin ↓ Book a call theseo-ai-crawlerlog-1.0.0.zip · 52 KB · version 1.0.0 · GPLv2 or later · no account, no registration, no key
Gianluca, who builds the tools and plugins on this site
image · whoever built the tool also explains it
FIG.01: What happens to a requestsheet 1/3

Three steps, and for an ordinary visitor it stops at step two

The plugin hooks into the requests that arrive at your server anyway. Nothing is added to your page, no script is loaded, and a visitor notices nothing. This is what happens, and nothing more.

requestserver side · no script in the page
step 1

A request comes in

WordPress receives a request for a page. The plugin looks at one thing: the user agent the other side sends along.

If the request has no user agent at all, it stops here.

step 2

Compared with the list

The user agent is compared with a list of 34 names built into the plugin itself, plus any names you've typed in yourself.

No match, and nothing further happens. That's the path for every ordinary visitor.

step 3

A row added

A match, and exactly one INSERT follows with four fields: the crawler's name, the page's address, the status code your server returned and the time.

The table has no other fields.

The name that goes into the table is the name from the list, not the string the other side sent. A crawler is recognised by what it claims to be, and anyone can claim anything. The log tells you what was claimed, not who it really was. This version has no reverse DNS check and no IP range check.
DOC.01

Your server log knows, you just don't read it

Every AI crawler that fetches your site leaves a trace in your web server's access log. In practice nobody looks. The file is huge, full of images and bot traffic, and most hosting packages keep it for only a few days. The question you actually have is small: did GPTBot come by, on which pages, and did my server return a 200 or a 404.

Since 2025 there's a second question that didn't used to exist

A crawler fetching your page can do three different things with it: put it in a search index, use it as a source while answering a question, or use it to train a model. Those are three separate forms of use, and until recently you had no way to say something different per form. Disallow was all or nothing.

This plugin does those two things side by side, because in practice it's a conversation

You see who's been by, and on that basis you decide what you allow.

DOC.02

What the plugin actually does

Two screens, and a third with the settings.

The log

One row every time a crawler from the list fetched a page on your site. Counted per crawler and per page, over 7, 30 or 90 days or over everything, with four totals at the top, a bar per day and an export to CSV of up to 20,000 rows.

The content signals

A screen where you state, per crawler group, whether you make your content available for search, as input for an AI answer, and for training. You see the exact rules that follow, spelled out on screen, before you save, and nothing is added to your robots.txt until you tick the box yourself. There's also a separate column that adds a genuine Disallow: / for a group. That's a different thing from a signal, it's off everywhere by default, and the screen keeps the two apart.

The settings

Measuring on or off, a retention period of 30, 90 or 365 days or keep everything, and a field for your own names, up to 25, one per line. Each line is searched for within the user agent, so MyBot also matches MyBot/1.0 (+https://example.com/bot).

FIG.02: What goes into robots.txtsheet 2/3

Three statements per group, and you see the rules before you save

A content signal is a statement about what you allow with your content once it's been fetched. There are three of them and they're independent of each other. Each signal can be yes, no, or no stated preference. No stated preference leaves the signal out of the line entirely, and that's exactly how you say you neither allow nor forbid that form of use. Everything starts on no stated preference, and the block is off.

/robots.txtexample · this is what the block looks like
# Content signals, written by AI Crawler Log # Explanation of the format: https://contentsignals.org/ User-agent: GPTBot Content-Signal: search=yes, ai-input=no, ai-train=no Allow: / User-agent: CCBot Content-Signal: ai-train=no Disallow: / # search and ai-input have no stated preference for CCBot, so they # don't appear in the line. The Disallow below is a separate choice and # is off everywhere until you switch it on yourself.
A signal is a statement of preference, not a technical block. It only works if the other party respects it. Standardisation is still under way in the IETF's AIPREF working group; the plugin writes the format the Content Signals Policy publishes and follows that text as it changes. If there's already a real robots.txt file at the root of your site, your web server serves that file and WordPress never gets involved. The screen checks for this and tells you.
DOC.04
WordPress6.0 or higher. Tested up to and including 7.0 in the plugin.
PHP7.4 or higher. The code has been checked against 8.5.
PermissionsYou need manage_options, so administrator.
Screen languageDutch on a Dutch site, English otherwise.
Also available on Full source code on GitHub The zip above and the source code on GitHub are the same version. This plugin isn't in the official WordPress plugin directory yet. It's ready for submission; until that's settled, the download above is the only place to get it. Checked on 19 August 2026.
STEP 1 · DOWNLOAD

Click the download button above. You get theseo-ai-crawlerlog-1.0.0.zip, 52 KB. Don't unzip it: WordPress wants the zip itself.

STEP 2 · UPLOAD

In your WordPress: Plugins, Add New Plugin, Upload Plugin. Choose the file and click Install Now.

STEP 3 · ACTIVATE

Click Activate Plugin. A new item appears in your admin menu. Your site doesn't change yet at this point.

STEP 4 · CONFIGURE

Open the settings screen and go through the options. Anything that touches your public site is off until you switch it on yourself.

DOC.03

What it doesn't do, and doesn't claim

This is here because it's the fastest way to get this topic wrong.

POINT 01

It doesn't tell you whether you're quoted

Whether ChatGPT, Perplexity or Gemini mentions your site in an answer can't be measured from your own server. A crawler visit isn't a citation, and the plugin doesn't pretend it is.
POINT 02

It doesn't verify the crawler

A user agent is a claim. Anyone posing as GPTBot ends up in the log as GPTBot. This version has no reverse DNS check and no IP range check, and the screen says so too.
POINT 03

Two names can never appear in the log

Google-Extended and Applebot-Extended only exist as a group in robots.txt and are never sent with a request. You can give them a signal, but they'll never show up in the list. The plugin says so on screen rather than silently counting zero.
POINT 04

An empty log usually isn't a fault

Three common causes: no crawler has been by yet, measuring is switched off, or a full-page cache or CDN answers the requests so they never reach PHP. In that last case almost no WordPress plugin can see it, unless it reads your CDN's log afterwards. The settings screen prints out the list of names so you can paste it into your cache plugin's exceptions.
FIG.03: What gets recorded per rowsheet 3/3

Four fields in, and six things deliberately left out

This is the whole table. There's no second table, no hidden column and no second place anything ends up. The right-hand column isn't a promise but the reason there's no personal data in this log that you'd need to write into a record of processing activities.

a row in the logthe whole table · four columns
is recorded
  • The crawler's name, as it appears in the plugin's list
  • The page's address, without a query string and without an anchor
  • The HTTP status code your server returned
  • The time, according to your server
is not recorded
  • The IP address, not in readable form and not as a fingerprint either
  • The full user agent string
  • A cookie or anything in localStorage or sessionStorage
  • A user number or a link to a WordPress account
  • The query string, so a key or email address in an address never ends up in it
  • Anything at all that leaves your server
The visitor's address is used exactly once: to count how many requests come in per minute, so a spoofed user agent can't flood the table. For that it becomes a key fingerprint inside a value that expires within a minute, and it never reaches the log table. Checked on 19 August 2026 in includes/class-theseo-crawlerlog-request.php.
DOC.05

Privacy: what is and isn't recorded

What goes into the table is set out above in FIG.03, field by field. What follows here is how you check it yourself, because a privacy paragraph you can't verify is just marketing copy.

POINT 01

No outgoing requests at all

Search the unzipped file for wp_remote_, curl_init, file_get_contents and fsockopen. Zero matches. There's no account, no key and no service behind it.
POINT 02

No IP address in the table

The address appears exactly once, in the rate limit, where it becomes a transient name via a key fingerprint that expires within a minute.
POINT 03

No cookie and no browser storage

The plugin loads no script at all on your front end. There's nothing that sets or reads anything on a visitor's device.
POINT 04

No query string

The address is rebuilt from scratch, so a session token or an email address someone put in an address never ends up in the table.
POINT 05

Retention period

A daily job deletes rows older than the set period, 90 days by default. Deactivating keeps the log and stops the job. Deleting drops the table, the settings and the job, readable in uninstall.php.

For the sake of completeness: this log is about machines, not people. It contains no data traceable to a person, which is why you can put it on a client site without needing anything added to a record of processing activities.

DOC.06

Frequently asked questions

Does the plugin slow down my site? The check runs on requests that carry a user agent, compares it with a list of short strings, and stops there for an ordinary visitor. Only a match produces an INSERT. Nothing is added to the page and no script is loaded.

Why isn't there a nonce on the logging?

Because logging doesn't accept input. It looks at requests that arrive on their own and writes a row server side. There's no address to send anything to. The admin screens, the export and every save action do use a nonce plus a permissions check.

Does it work with a cache plugin?

Only to the extent that requests reach PHP. If your cache or your CDN answers the request itself, WordPress never sees it, and so neither does this plugin. The settings screen prints out the list of names so you can exclude them from the cache.

Does it work on multisite?

Yes. Every site in the network gets its own table, its own settings and its own robots.txt block. Deleting cleans up every site.

Can I put it on client sites?

Yes. The data stays within that client's WordPress, and the settings screen states plainly what's kept. There's no personal data to show a data protection officer.

Does it change my robots.txt as soon as I activate it?

No. Logging starts straight away; the robots.txt side is off. Nothing is added until you set the signals, review the preview and tick the box.

DOC.07 · Next step

What you get out of the first week

Turn it on for a week and then see who's been by. That conversation is almost always shorter and more concrete than the conversation about AI visibility in general. Want to know afterwards how to make those pages citable too, not just fetchable? The AI snippet previewer is the next step, and the llms.txt generator states in plain language what your site is. Want to talk it through, book a call.

Section · Next step24/7 available · sales@theseo.nl
Book a call ↗
Written by GianlucaFounder. Building visibility for Dutch businesses since 2017, in Google and in AI answers. More about the institute.