The Kirkpatrick Evaluation Coach skill for Claude
You give your training goal and the behaviour change you want to see, and you get back a measurement plan across all four of Kirkpatrick's levels: a satisfaction survey, criteria for a pre- and post-test, an observation schedule for day 30, 60 and 90 with behavioural indicators for line managers, and the link to three to five existing business KPIs. Every figure comes from you or appears as [FILL IN] in the plan, with a suggestion for how to get it. Target scores and ROI percentages are not things the skill invents.
the rules above come from the zip on this page · SKILL.md is 9,828 bytes
Training goal in, measurement plan out
reg. K.001This is the worked example that appears in the SKILL.md itself, shown here in shortened form. Note the amounts and the target values. They are there, but never as a claim by the skill: the client sets the target score, the minimum improvement gets confirmed, and where there is no baseline measurement it simply says [FILL IN].
What the Kirkpatrick Evaluation Coach is
The Kirkpatrick Evaluation Coach is a free skill from our own skill library for Claude. A skill is an instruction file, SKILL.md, that gives an AI assistant a fixed way of working for one task. No software to install, no subscription, no link to your LMS: a single 1,390-word text file that tells Claude how Kirkpatrick's four levels work, what he needs to ask you, in what order the measurement plan gets built, and which figures he must never invent himself. You download the zip at the top of this page, drop it into your AI environment, and from that moment Claude builds measurement plans following this logic.
The problem the skill solves is named in the file itself and will be familiar to anyone who has ever organised a training course. A satisfaction survey tells you how people felt. Level 3 and 4 tell you whether anything actually changed. Most evaluations get stuck at level 1, leaving you with enthusiastic participants and a board asking what it delivered. This skill carries the measurement plan through to the figures that board actually looks at, without pretending that this proves the training was the cause.
It is written for training courses, onboarding, compliance programmes and AI adoption programmes. No prior knowledge of Kirkpatrick is needed: the levels, the design logic and the pitfalls all sit inside the file. It is part of the skill library NL that we make freely available through our AI and automation service, with no account and no follow-up sales email. New to skills? What are Claude skills NL explains first exactly what you are downloading.
One thing you should know upfront, and the file is strict about it: this is a measurement framework, not an instructional design and not a scientific effectiveness study. The measurement plan makes the effect plausible and discussable. Without a control group, level 4 shows correlation and not cause, and according to the skill that sentence belongs in your reporting too.
Why evaluations get stuck at level 1
Almost every training course ends with a form. Was it useful, was the trainer good, how were the facilities. That form is easy to send out, almost always comes back with a decent score, and says precisely nothing about what happens on the shop floor on Monday. Four causes keep that pattern in place, and the skill tackles all four.
The measurement is only thought up afterwards. By then the training is already designed, the participants are already chosen and the moment for a baseline measurement has passed. The New World Kirkpatrick Model reverses that: you start at level 4, the desired business result, and work back to what people need to do, know and experience to get there. The skill follows that order literally: first the question of which result the training should affect, and only then does it build level 1.
There is no nameable behaviour. Without a concrete description of what people do differently after the training, there is nothing to observe. That is why the skill asks that question first, and why it does not proceed without an answer. The file says it plainly: no nameable behaviour, no measurement plan.
Indicators are feelings, not behaviour. Shows engagement is not an indicator, because two line managers will see something different in it. Shares a method with the team every week is, because it either happened or it did not. It refuses to deliver observation schedules without observable behaviour, and that difference is exactly where most behaviour measurement comes unstuck.
And level 4 gets either skipped or overstated. Skipped because it is hard, or a KPI rise gets attributed entirely to the training. Neither happens here, and that is a deliberate choice. It links three to five existing KPIs with a data source and measurement frequency, and names in the same breath which other factors move the same KPI.
In diff form, with indicators from the skill's example. On the left what usually ends up in an observation form, on the right what you can actually tick off.
Designed from the top down
reg. K.002The four levels as they appear in the SKILL.md, but in the order in which the skill designs them. You measure from the bottom up, from reaction to results. You design the other way round: you start at the business result and work back. The highlighted row at the top is therefore the skill's starting point, not its final stop.
What is actually in the SKILL.md
A skill is only as good as its instructions, so we simply describe them here. The file opens with frontmatter stating when Claude should pick up the skill. Not just for the obvious questions like kirkpatrick, evaluating training, measurement plan for a training course or justifying a training budget, but also for complaints: I don't know whether my training delivers anything, we only send out a satisfaction form, everyone was enthusiastic but nobody does anything differently, after the training everything sinks back in, my evaluations only tell me something about the sandwiches. Say something like that to Claude while the skill is loaded, and it picks it up automatically.
Next comes the theory the skill rests on, written out in the file itself. Donald Kirkpatrick described his four evaluation levels in 1959 in a series of articles in the Journal of the American Society of Training Directors, drawn from his doctoral dissertation. In 1994 he brought the model together in Evaluating Training Programs: The Four Levels. His son James and daughter-in-law Wendy updated it in 2016 as the New World Kirkpatrick Model, with more emphasis on designing the measurement upfront and on support after the training.
The core of the file is a list of eight fixed steps. First it asks targeted questions: the training goal, the target group and group size, duration and format, and above all the desired behaviour change on the shop floor. It works back from level 4 and looks for three to five existing KPIs. It builds level 1 as a short survey on relevance, applicability and engagement, with a target score that you set. Level 2 becomes a set of test criteria for a pre- and post-test that match the desired behaviour.
It builds level 3 as an observation schedule for day 30, 60 and 90, with three to five observable indicators per measurement point and a form to fill in. Level 4 becomes the link between those indicators and the KPIs, with measurement frequency, data source and the caveat about which other factors move alongside it. It states the resources needed honestly: a survey tool, a spreadsheet and an observation form are enough. And it delivers a timeline from day 0 to day 90, with the owner of the measurement at each point.
Before anything gets built, the skill runs through a mandatory input checklist: the training goal in one sentence and the desired behaviour change, the target group and number of participants plus whether their line managers are willing to help with observation, duration and format, the existing business KPIs with current values if you have them, the available tools, and the timeframe in which the organisation wants to see results plus who needs to approve the measurement plan. It only asks about what is missing.
The output follows a fixed order of six parts: goal and behaviour change with the assumptions you need to confirm, then the four levels each with their measurement point and owner, and finally the timeline with every measurement point laid out in a row. Want to learn to write files like this yourself? Writing SKILL.md NL explains the structure.
Day 0 to day 90
reg. K.003The sixth block of the output is a timeline of every measurement point with an owner for each. That block is the reason a measurement plan stays alive: without a name against a date, the measurement on day 60 does not happen. The timeline from the skill's example, with the role each measurement point plays according to that same example.
The theory the skill rests on
The SKILL.md names five sources by name, and that is deliberate: you can read them yourself and decide whether you agree with the choices. Notably, one of them is a critical source. That is not carelessness, that is why the skill words itself so carefully.
The model comes from Donald Kirkpatrick. He described the four levels in 1959 in a series of articles in the Journal of the American Society of Training Directors, drawn from his doctoral dissertation, and brought it together in 1994 in Evaluating Training Programs: The Four Levels. That is still the skeleton of every measurement plan the skill produces.
The design logic comes from James and Wendy Kirkpatrick. Their New World Kirkpatrick Model from 2016 places emphasis on two things that were underexposed in the original version: designing the measurement upfront instead of thinking it up afterwards, and support after the training. That is where the reversed order in FIG.02 comes from.
The fifth level comes from Jack Phillips. Return on Investment in Training and Performance Improvement Programs from 1997 added ROI on top of level 4, with the accompanying caveats. The skill names that source, but it does not calculate an ROI and does not invent a single percentage. If you want to work out the amounts behind that kind of calculation, you are better off with the Pricing Strategy Coach skill NL, which deals with exactly that value question.
The criticism comes from Alliger and Janak. Their article Kirkpatrick's Levels of Training Criteria: Thirty Years Later from 1989 in Personnel Psychology is about the assumed causality between the levels: the idea that satisfied participants also learn, and that whoever learns also applies it, turns out not to be a given. That source is in the file and explains why the skill never presents level 1 as a predictor of level 2.
Even without installing the skill you can improve your own evaluation with these principles: start with the result you want to see, write down behaviour you can tick off, agree an owner per measurement point, and put in your reporting whatever else was putting pressure on that KPI.
What the skill refuses
reg. K.004The SKILL.md contains a list of seven things the skill never does, and in evaluation work that list is half the product. A measurement plan with invented target figures looks professional and is worthless. In conversation those rules play out like this: each row is a request you might make, with the response the skill gives according to its own instructions.
Installing in Claude Code, Claude.ai or Codex
The zip contains one folder, ai-kirkpatrick-evaluatiecoach, with the SKILL.md inside. Installing is simply a matter of putting that file in the right place, and that place differs per environment. SKILL.md has been an open standard since December 2025, so the same skill also works in Codex, Cursor and Gemini CLI. So you are not downloading a Claude file but a working instruction that any modern AI assistant can read.
- Unpack the zip into
~/.claude/skills/(or.claude/skills/in your project). - Claude then recognises the skill automatically as soon as you start talking about evaluating training, measurement plans or Kirkpatrick.
- You can also call it directly, with
/ai-kirkpatrick-evaluatiecoach.
- Go to Customize and then Skills.
- Upload the zip there as a skill.
- Or paste the contents of SKILL.md into a Project's project instructions.
- Open
AGENTS.mdin your repo. - Paste the contents of SKILL.md into it, or place SKILL.md alongside it as a separate file and refer to it from
AGENTS.md. - Codex reads that in with every session.
After that, using it is simple: describe your training goal and, above all, what people need to do differently afterwards. If you get stuck unpacking the zip or finding the right folder, Installing Claude skills NL walks through it step by step. And if you want the broader explanation of working with AI first, you will find it in the knowledge base.
When you do and don't use it
It is strongest for programmes where someone will later ask what it delivered: an AI adoption programme, an onboarding programme, a compliance programme, a training series with a budget around it. The earlier you deploy it, the more you get out of it, because the baseline measurement on day 0 is the one measurement point you cannot catch up on later. It is also useful if you just want to sharpen up what you actually want to see change: that first question about behaviour often opens up a different conversation with the client.
There are also situations where you are better off leaving it aside, and the file names them itself. Designing the training is one of them: this is a measurement framework, not an instructional design. What happens in the training and how you build up the material is work for a different kind of help, for example the Bloom Taxonomy Coach skill, which writes learning objectives and test questions per thinking level. Statistical proof is the second limit: the measurement plan makes the effect plausible and discussable, it is not a scientific effectiveness study with a control group. And the third is assessing people: the observation scores are there to evaluate the training, not to fill out a personnel file.
One more honest limit the skill names itself: with small groups and short timeframes, level 4 is often hard to prove. That does not mean you skip it, but that the reporting honestly states what you can and cannot see. If you are working towards a decision on whether to continue or stop, that distinction matters more than the figure itself.
Run it yourself or have it run
reg. K.005This skill is the free do-it-yourself version of work we also deliver as a service. The skill stays complete and free of catches, but be aware what a skill is: it teaches your AI how to do something, while every new session starts empty. The skill is not the engine and not the memory. You prompt, you supply the context again every time, you check. For a measurement plan running over ninety days that is a real drawback, because the measurement on day 60 does not come back onto your screen by itself.
where you are now The skill: you are the engine You run the Kirkpatrick Evaluation Coach yourself in Claude, Codex or Cursor. Costs nothing, works today, and you keep it fully in your own hands: no trial period, no locked-off parts. The limit is your own time: it only happens when you prompt, and you explain your KPIs and your organisation afresh every session.
have it prepared The Reporting employee: it is ready without you prompting Exactly this work, but as a service: the Reporting employee prepares weekly and monthly reports without you having to sit down behind Claude for it, and that is exactly what a measurement plan needs on day 30, 60 and 90. Control stays with you, because output stays a draft until a person approves it. We deliver this through Mansotti, the company of which TheSEO is the trading name, which alongside the Reporting employee also builds a Quote employee and a Sales employee (prospect research). Read what an AI employee is and does.
everything from one source Jarvis: all your AIs work from the same company knowledge The skill teaches the AI, the brain is where the memory lives. If you want all your AIs to work from the same company knowledge: that is Jarvis, the organisation brain. It connects ChatGPT, Claude, Codex and your people to the same projects, core knowledge and decisions, so your next session does not start over. This skill benefits too, because you no longer need to supply your KPI definitions, your baseline measurements and the agreement on who measures what, session after session. What that delivers in practice, from the plans to your first week, you can read at Jarvis itself.
What Jarvis actually delivers
reg. K.006Step 3 deserves more than a paragraph, because this is the difference between a clever chat and a system you can build on. Jarvis is the organisation brain: it remembers what your AIs need to know, divides the work and keeps track of what happened. For a ninety-day measurement plan that is not a luxury but the core of it, because a measurement nobody retrieves is not a measurement.
We have been running on this system ourselves for months. Every agent session, every task and every decision is logged in it and can be read back. A new session therefore does not start blank: it first fetches the recorded decisions, the running projects and the latest changes, and carries on where the previous one stopped. So we are not describing a promise but the way we work every day.
See the four plans at jarvis/pricing NL. Through the waiting list NL you only pass on your preferred plan, without obligation. That does not yet create an account, order or payment obligation. Business bespoke work we discuss first.
The skills around it
reg. K.007A measurement plan never stands alone: a training course comes before it and a report follows after. These skills from the same library each cover a different piece of that chain.
Before the training
the designThe measurement plan measures what the training aims for. These two cover that first part.
Bloom Taxonomy CoachWrites learning objectives and test questions per thinking level, exactly what level 2 needs.SKILL Feynman Technique CoachKeeps simplifying material until the gaps in understanding become visible.SKILLAround the measurement
the figuresLevel 4 leans on KPIs that already exist. These two are about those figures themselves.
Balanced Scorecard CoachSorts your business KPIs across four perspectives, so you know which ones to link.SKILL OKR CoachTurns goals into measurable key results, with the same rule: no figure without a source.SKILLAfter the training
the behaviourLevel 3 is about what people keep doing. These two are about that staying power.
Growth Mindset CoachFor the participant who slips back into the old way of working after two weeks.SKILL Deliberate Practice CoachSets up targeted practice with feedback, so behaviour does not fade after day 30.SKILLLooking further
the contextWhere this skill comes from and what else there is.
The whole skill library100 free skills in a row, sorted by topic.HUB NL AI and automationThe service behind it: from loose skills to working automation.SRV AI trainingThe training this measurement plan fits around, if your team is getting started with AI.SRV NL Knowledge baseArticles on SEO, AI and online visibility, searchable.DOCFrequently asked questions
What does the Kirkpatrick Evaluation Coach skill cost?
Nothing. The skill is free, comes under the MIT licence, and you don't need to create an account or leave an email address. You download a 4.1 KB zip containing a folder and a single file, SKILL.md, and that is the complete skill. There is no paid version and no sales email follows afterwards.
Does this skill also work in Codex, Cursor or Gemini CLI?
Yes. SKILL.md has been an open standard since December 2025, so the same file also works in Codex, Cursor, Gemini CLI and other tools that follow the standard. In Codex you unpack the zip into .agents/skills/ in your project, or into ~/.agents/skills/ for all your projects; Codex has supported SKILL.md directly since the open standard of December 2025. Putting the contents of SKILL.md into your AGENTS.md still works too. The instructions are just readable text, so any assistant that accepts instruction files can work with it.
Does the skill calculate the ROI of my training?
No. Inventing ROI percentages is on the list of things the skill never does. The fifth level, ROI, comes from Jack Phillips's 1997 work and is in the sources list, with its caveats included. What the skill does do is build level 4: linking three to five existing business KPIs to the behavioural indicators, with data source, measurement frequency and the external factors that influence the same KPI.
Does the skill invent target scores and baseline measurements if I don't have any?
No. Every figure comes from you or becomes [FILL IN] with a suggestion for how to get it. In the worked example, a target score of 4.0 out of 5 is stated explicitly as something the client sets, and a minimum improvement of 25 per cent as something the client confirms. If a baseline measurement is missing, it says [FILL IN: baseline measurement needed].
Does a measurement plan prove that my training works?
No, and the skill says so itself. Without a control group, level 4 shows correlation and not a causal link, and that sentence belongs in the reporting too. The model has further limits: the levels are not a proven causal chain, satisfaction is a poor predictor of learning outcomes, and with small groups and short timeframes level 4 is often hard to prove.
Can I measure only level 1 and 2?
Yes, that is possible, but not silently. The skill does not skip level 3 and 4 because they are hard. If you deliberately choose only reaction and learning, it names what stays unanswered as a result: whether people actually apply what they learned, and whether anything changes in the figures the board looks at.
Which tools do I need to run the measurement plan?
According to the file, a survey tool, a spreadsheet and an observation form are enough. Heavy measurement systems are not needed, and the skill states that explicitly as part of the plan. In the input checklist it does ask which tools you already have, so the measurement plan fits what you have rather than what you would need to buy.
Behaviour first, then the figures
The hardest question in this whole framework is the first one: what do people demonstrably do differently after the training. Anyone with a clear answer to that has the measurement plan half done, and usually a better training too. Want your team to get going with AI and be able to show straight away what that changes? That is exactly the combination we deliver in our AI training NL. And if you want to talk further about what else AI can take on in your organisation, from loose skills to full automation, we simply do that in a conversation.