The AI First Principles Thinker skill for Claude
You put forward a plan and Claude strips it down to what is demonstrably true. What remains falls into three piles: the things that never change, the things you can change through contract or regulation, and the conventions that pass themselves off as fact. From that third pile come eight to twelve assumptions, each written as a claim that could turn out to be wrong, each scored on probability and impact. The two to four most dangerous get a cheap test you can run this week. This is first principles thinking in practice.
the lines above come from the zip on this page · SKILL.md is 17,104 bytes
A plan split into three kinds of statement
reg. E.001The case below appears verbatim in the SKILL.md, shown here in shortened form. Notice what happens first: the goal is rewritten into a measurable sentence with a deadline, and the word launch is removed because a launch is an event, not a goal. Only after that does the plan get split apart. The third pile is almost always the longest, and that is where the work lies.
# removed from the wording: launch. A launch is an event, not a goal.
# paying participants is a goal, because that is countable
What first principles thinking is
First principles thinking is a way of reasoning where you reduce a question to the things you can prove, and only then build it back up. You do not start from how it has always been done, from what the competitor does, or from what feels logical, but from the bare facts that remain once you remove every assumption. It is the same tool people mean by thinking from the ground up, or getting back to first principles.
The contrast that makes the method so useful sits between two ways of reasoning. The first is reasoning by analogy: you look at what others do and adapt it. Fast, safe, and it never delivers more than a variation on what already exists. The second is reasoning from first principles: you break the problem down to what is demonstrably true and build back up from there. Slow, uncomfortable, and the only route to an answer that departs from the norm. That is exactly how the skill file puts it too, and that phrasing is sharper than the usual description because it states the price of the method straight away.
The AI First Principles Thinker translates that way of thinking into a fixed method for Claude. A skill is an instruction file, SKILL.md, that gives an AI assistant a fixed approach for a task. No software, no subscription, no connection to your systems: a text file of 2,744 words that tells Claude how to strip a plan down, how many assumptions to surface, how to score them, what output to deliver, and what it must never make up under any circumstances. You download the zip at the top of this page, put it in your Claude environment, and from that moment every plan you put forward goes through this route.
What sets the skill apart from just thinking critically along with you comes down to three things. First, when it breaks a plan down it separates three kinds of statement: laws of nature and mathematical truths that never change, contractual and legal facts that can change but not because you want them to, and conventions. That third category is where the gain lies, because conventions present themselves as facts. Second, it writes every assumption so that it could turn out to be wrong. Not that the market is big enough, but a claim with a number and an amount in it that you can check. Third, it keeps probability and impact strictly apart, each as its own figure from 1 to 5, because a merged gut-feel score does not point you anywhere.
This file belongs to our skill library NL, where 100 skills are free to download, and comes out of our AI and automation service. If you are working with this kind of file for the first time, start with what Claude skills are NL: it explains in plain language what a skill can and cannot do.
Why you cannot see your own assumptions
An assumption you recognise is not a problem, because then you can test it. The problem lies in the assumptions that pass themselves off as fact. There are three ways that happens, and the skill has a countermove for each.
The convention that passes itself off as a law of nature. That is how we do it here, that is how the market works, that is how it is supposed to be. This is the richest layer, because conventions are invisible on purpose: they only work as long as nobody notices them any more. The breakdown block in the output forces the separation by putting three lists one below the other, and the file adds, drily, that the third list is usually the longest and the most interesting. In the case study that comes out as something like: that a course has to cost €495 because others charge that. Put that way, it is no longer pricing policy but an assumption with a price tag on it.
The claim that could never turn out to be wrong. The market is big enough. Our target audience needs this. Sentences like that feel like analysis but are not, because there is no conceivable outcome that would disprove them. The skill's writing rules forbid them and immediately give the alternative: write a claim with a number and an amount in it instead, so that you can check it. In review mode that is a separate check, the falsifiability check, which rewrites every unfalsifiable assumption into a claim that can be tested.
The judgement that folds probability and impact into one. This is a big risk: sounds like a score, points you nowhere. You do not know whether it is big because it is likely to go wrong or because it is expensive if it does, even though you act very differently on those two. With a high probability and low impact you run a quick, cheap test. With a low probability and high impact you build a safety net. That is why there are always two separate figures, and merging them into a gut-feel score is explicitly on the list of things it never does.
There is a fourth thing underneath all three, and it is stated right in the opening lines of the file: the skill confirms nothing out of politeness. That is not a matter of tone but a working instruction. The assumption the whole plan rests on, and that nobody has checked, is usually exactly the assumption everyone in the industry shares, and it survives every conversation in which the participants want to be nice to each other.
In diff form, with the same claims first as they usually appear in a plan and then as the skill rewrites them:
Probability times impact, two separate figures
reg. E.002Block 4 of the output. Every assumption gets a probability of being wrong from 1 to 5, an impact if it turns out to be wrong from 1 to 5, and the product sets the order. Alongside that sits the evidence strength: proven, plausible or unproven. The top rows become the primary risks. The four rows below are the top four from the example in the SKILL.md, with the figures exactly as they appear there.
What is actually in the SKILL.md
A skill is only as good as its instructions, so we simply describe them here. The file runs to 2,744 words and consists of ten chapters: the theory, what the skill always does, a mandatory input checklist, the writing rules, the output structure, a worked example, when to use which variant, how to review existing work, what the skill never does, and the sources.
It opens with frontmatter stating when Claude should pick up the skill. Not just for the explicit terms like first principles thinking, first principles, thinking from the ground up, surfacing assumptions, assumption mapping, riskiest assumption and breaking down a cost structure. Also for questions that have nothing to do with a framework: does my business plan hold up, what am I missing, why is this so expensive, can this be fundamentally different.
And for exasperated remarks you would sooner hear in a canteen than in a strategy session: everyone in the industry works this way so we do too, we copied this from a competitor, my investor asks why this works and I do not have a good answer, our price is based on what everyone else charges.
Right below that sits the role description, and it is unusually strict for an AI instruction: the skill confirms nothing out of politeness, hunts for the assumption the whole plan rests on that nobody has checked because everyone in the industry shares it, and never makes up figures, market data or research findings.
The core is a list of eight fixed actions. It rewrites the goal into a measurable sentence with a deadline and strips out the solution if one is already hidden inside it. It strips the plan down and separates laws of nature and mathematical facts from contractual facts and from conventions. It surfaces eight to twelve assumptions, spread across the customer, the market, execution, technology, money and regulation, each as a falsifiable claim. It scores every assumption on probability and impact, from 1 to 5, and records the evidence strength.
After that it gets sharper. It flags the highest-scoring ones as primary risks, usually two to four. For each risk it delivers a redesign step you can carry out this week plus the cheapest test. It rebuilds the plan from what remained and states for each step whether it departs from the original plan and why. And it closes with the risks you consciously accept, so that there is a choice in place of a blind spot.
Before anything happens, the skill runs through a mandatory input checklist of seven points: the plan in your own words, the deadline by which it has to succeed, what is at stake in money, time, reputation or business, where the plan came from, the hard constraints, what you already know for certain and what that is based on, and what has already been tried in this direction. It only asks about what is missing and then stops asking, so you do not get a form, just a handful of targeted questions at most.
The output is fixed into eight blocks in a set order: the goal sharpened, broken down into three lists, the assumptions, the score table, the primary risks, the redesign per risk, the plan rebuilt in five to eight steps, and consciously accepted. The file also contains a worked case study about an online course, six situations where you pick which variant, a review mode with six checks for an existing plan or an existing risk analysis, seven things the skill never does, and six sources. If you want to see how a file like this is technically put together, from frontmatter to refusal list, writing a SKILL.md NL explains that step by step.
A striking number of rules are about language, and they are not all cosmetic. Clear Dutch, addressed as you, no dashes in running text, AI in capitals, no lists of exactly three, concrete over abstract, a figure over an adjective. But the last writing rule is substantive and maybe the most important in the whole file: word every assumption so that it could turn out to be wrong. That one rule decides whether the rest of the analysis is about anything at all.
For every risk, a test and an adjustment
reg. E.003Block 6 is where most analyses stop and this skill keeps going. For every primary risk it delivers two different things: the cheapest test that shows whether the assumption holds, and an adjustment to the plan that shrinks the risk even if the assumption does hold. That distinction is the point. A test gives you information, an adjustment gives you a plan that depends less on that information.
The theory the skill rests on
The SKILL.md closes with six named sources. That is deliberate: it lets you look them up and decide for yourself whether you agree. Each source explains a choice in the method.
The idea comes from Aristotle. He wrote about first principles: the starting points you cannot derive from anything else and on which all further reasoning rests, among other places in the Metaphysics and the Posterior Analytics. The term first principles traces back to that work. That is exactly what the breakdown block does. Not searching for the best explanation, but for the layer that needs no explanation at all.
The method comes from Descartes. His Discours de la méthode from 1637 made methodical doubt the road to certain knowledge: set aside everything that can be doubted and only then build back up. The two-part break-down-and-build-up you see in FIG.01 and FIG.03 comes from here.
The scoring comes from Bland and Osterwalder. Testing Business Ideas from 2019 worked out mapping and testing assumptions for entrepreneurs, in the form of an assumption map that plots importance against the amount of evidence. That is where the idea comes from that an assumption has two independent properties and that you must not fold them together. If you then want to actually run and measure those tests, the Mom Test Coach skill connects directly to that: it is about customer conversations in which people do not just tell you what you want to hear.
The premortem comes from Gary Klein. His article Performing a Project Premortem in Harvard Business Review from 2007 describes the exercise in which you project yourself forward to a moment when the plan has failed and write down the story of that failure. It is named as a variant and comes with a practical rule for teams: have everyone write separately before you bring them together, otherwise everyone follows the first speaker. As a stand-alone skill, that exercise lives at the Pre-Mortem Analyst NL.
Inversion comes from Jacobi via Munger. Charlie Munger popularised inversion as a thinking tool, collected in Poor Charlie's Almanack from 2005, and the rule itself is generally attributed to the mathematician Carl Jacobi. You do not ask how this succeeds but how this is guaranteed to fail, and then avoid that route. That one also has its own file: the Inversion Thinker.
Its business fame comes from Elon Musk. He used the approach to reduce the cost structure of rockets and batteries down to the price of the raw materials, instead of starting from what such parts had historically cost. The file is careful about this: that application became widely known through public interviews and cannot be attributed to a publication. In the skill the same move is called the cost breakdown, and you can turn it loose on any price by breaking it into raw material, labour, transport and margin, and asking of each component why it is what it is and who sets that.
Even without installing the skill, you can get further with these principles: split your plan into three piles, rewrite every claim into something that could turn out to be wrong, give probability and impact two separate figures, and for the top two work out the cheapest way to find out.
What the skill refuses
reg. E.004The SKILL.md contains a list of seven things the skill never does, and that list matters at least as much as what it does do. A risk analysis nobody dares to act on is wasted time. In conversation those rules play out like this: each rule is a request you might make, with the response the skill gives according to its own instructions.
Installing in Claude Code, Claude.ai or Codex
The zip contains the folder eerste-principes-denker with the SKILL.md inside it. Installing is a matter of putting the file in the right place, and that place differs per environment. SKILL.md has been an open standard since December 2025, so the same skill also works in Codex, Cursor and Gemini CLI. So you are not downloading a Claude file but a working instruction that any modern AI assistant can read.
- Unzip it into
~/.claude/skills/(or.claude/skills/in your project). - Claude then recognises the skill automatically as soon as you put forward a plan or start talking about assumptions.
- You can also call it directly, with
/ai-eerste-principes-denker.
- Go to Customize and then Skills.
- Upload the zip there as a skill.
- Or paste the contents of SKILL.md into the project instructions of a Project.
- Open
AGENTS.mdin your repo. - Paste the contents of SKILL.md into it, or place SKILL.md alongside it as a separate file and refer to it from
AGENTS.md. - Codex reads that along at every session.
After that, using it is simple: describe your plan in your own words, with the deadline attached and what is at stake. You do not need to phrase it neatly, because it is precisely the messy phrasing that gives away the conventions. It runs through the input checklist, only asks about what is missing, and then works through its eight steps. If the installation gets stuck anywhere, the full explanation per environment is at installing Claude skills NL, including the mistakes everyone makes with folders and permissions. And if you want to learn to work more broadly with AI, you will find the background articles in the knowledge base.
When to use it and when not to
There is a dedicated chapter on choosing a variant, and it is more useful than general advice. Use the full breakdown with assumption scoring on plans where something real is at stake: an investment, a product launch, a strategic pivot. If the question is about price or margin, choose the cost breakdown and split the price into components, asking of each one why it is that high and who sets that. If the plan is already fixed and you mainly want to know what could go wrong, inversion is faster: ask how you would guarantee this fails and read the answers back as a risk list.
If a team is involved, the premortem works better than a conversation, with a practical rule attached that saves a lot of sessions: have everyone write the failure story separately before you bring them together, otherwise everyone follows the first speaker. And if a single claim carries the whole argument, socratic questioning is enough: keep asking how you know that, until you land on evidence or on thin air.
There are also situations where you are better off leaving it alone, and unlike many skills, that is stated in the file itself. Skip it for low-cost routine jobs, because thinking from the ground up costs time, and you only earn that time back on decisions that are hard to reverse. And do not use it if you just want a decision confirmed, because the skill confirms nothing out of politeness. That is not a minor caveat: if the decision has effectively already been made and the analysis is only there to justify it, this skill mainly produces friction.
Besides drafting, it can also assess. Hand it an existing plan, a business case or a risk analysis, and you get six checks back. The convention check flags every sentence that presents a choice as a given, with sentences starting with that is just how it is or in our industry as prime candidates. The analogy check names what has been copied from someone else and what the evidence is that it works there.
The other four checks are about the assumptions themselves. The falsifiability check rewrites every assumption that could not turn out to be wrong. The completeness check looks at which category is missing, and the file predicts which two that usually are: regulation and the assumptions about the customer. The scoring check looks at whether probability and impact were scored separately or folded into a judgement that points nowhere. And after that comes a rebuilt version following the standard structure.
One more boundary worth naming, and one the file does not state: the breakdown produces clarity, not certainty. After the score table you know which two to four claims your plan rests on, and that is progress, because before you did not know that either, you just thought you did. But you still have to test those claims afterwards, and that costs time or money. Anyone who uses the outcome to tear everything up without running a single test has turned the method on its head. If you want to go from a problem that has already occurred down to the cause underneath it, the 5 Whys Analyst skill is the tool for that, not this one.
Run it yourself, or have it run
reg. E.005This skill is the free do-it-yourself version of work we also deliver as a service. It stays complete and with no catches. Do be aware of what a skill is: it teaches your AI how to do something, while every new session starts empty. So a skill is neither the engine nor, certainly, the memory. You prompt, you supply the context again every time, you check the result. Anyone who wants it differently has two further steps: hand over the engine, or sort out the memory.
where you are now The skill: you are the engine You run the AI First Principles Thinker yourself in Claude, Codex or Cursor. Costs nothing, works today, and you keep it entirely in your own hands: no trial period, no locked-off parts. The limit is your own time: the breakdown only happens when you start it, and you have to supply last month's score table again yourself.
only if it recurs The AI employee: usually not the next step here None of our roles fit here, and we would rather say that than sell it to you anyway. The three roles we set up ready-made are the Quote Employee (sorting incoming requests and preparing draft quotes), the Sales Employee (prospect research and outreach drafts) and the Reporting Employee (summaries and weekly and monthly reports from your own data). What this skill delivers, surfacing assumptions and scoring them on probability and impact, fits none of the three. What can be done is a custom-built role through Mansotti, the company of which TheSEO is the trading name, but only if the tests that come out of it are tracked and ticked off somewhere. If this work stays a single session a year for you, skip step 2: being able to look up last month's score table is worth more than a service, and that is step 3. What an AI employee actually does is on that page.
everything from one source Jarvis: all your AIs work from the same company knowledge A skill teaches the AI something; in the brain that sticks. If all your AIs need to work from the same company knowledge, that is Jarvis, the organisation brain. It connects ChatGPT, Claude, Codex and your people to the same projects, core knowledge and decisions, so your next AI session does not start from scratch. With this skill that is extra noticeable, because an assumption you tested last month ought to stay proven afterwards, not go back to unproven. Client knowledge stays isolated and every step leaves a verifiable trail. What that delivers in concrete terms, from the plans to your first week, you can read at Jarvis itself.
What Jarvis concretely delivers
reg. E.006Step 3 deserves more than a paragraph, because this is the difference between a clever chat and a system you can build on. Jarvis is the organisation brain: it remembers what your AIs need to know, divides up the work and keeps track of what happened. For this skill that matters a great deal, because a score table is only useful once the evidence strength moves along with what you have tested. You notice it first at the start of a new session.
We have been running our own business on this system for months. Every agent session, every task and every decision is logged in it and can be read back. A new session therefore does not start blank: it first retrieves the recorded decisions, the running projects and the latest changes, and carries on where the previous one stopped. So not a promise on a sales page, but the way of working we ourselves are in every day.
See the four plans at jarvis/pricing NL. Via the waiting list NL you only indicate your preferred plan, without obligation. That does not yet create an account, an order or a payment obligation. We discuss business custom work first.
The skills around it
reg. E.007The variants chapter of this SKILL.md names four other thinking tools, and three of them have their own file in the same library. Below they are, together with the skills that come before and after.
The variants from the file
same familyIt names these itself as an alternative route when the context calls for it.
Inversion ThinkerNot how this succeeds, but how this is guaranteed to fail. Attributed to Jacobi, popularised by Munger.SKILL Pre-Mortem AnalystThe failure story from a year from now, written in advance. Works with a team, provided everyone writes separately.SKILL NL Socratic Method CoachFor every claim, the question of how you know that, until you land on evidence or on thin air.SKILL NL Pricing Strategy CoachFor when the cost breakdown shows that your price is a convention and not a calculation.SKILL NLBefore you break it down
diagnosisSometimes the question is not what foundation your plan has, but why something went wrong.
5 Whys AnalystDigs five layers through a problem that has already occurred, with a certainty label at every layer.SKILL Pareto AnalystLooks for the few causes that explain most of the problem.SKILL MECE Problem Structure CoachSplits a question into boxes that do not overlap and together cover everything.SKILL Feynman Technique CoachHas you explain it until a child would understand. The gaps in your explanation are your unproven assumptions.SKILLAfter you have scored it
testing and attackingA score table is only worth something once the top rows are actually tested.
Mom Test CoachCustomer conversations where people do not just tell you what you want to hear. The cheapest test for every customer assumption.SKILL Red Team AnalystAttacks the rebuilt plan from the opposing side and scores the assumptions on damage.SKILL Cognitive Bias DetectorFlags the thinking errors that make you write off an unproven assumption as proven.SKILL Bayesian Thinking CoachFor when an assumption becomes not a yes or no but a probability you adjust as evidence comes in.SKILLThe basics first
the explainerNever worked with skills before? Start here, and the file falls into place quickly.
What are Claude skillsIn plain language: what a SKILL.md is, what it can and cannot do, and why it is not software.BLOG NL Installing Claude skillsStep by step per environment, including the mistakes everyone makes with folders and permissions.BLOG NL Writing a SKILL.md yourselfHow the structure works, from frontmatter to refusal list, if you want to make one yourself.BLOG NL The whole skill libraryAll 100 free skills in one place, sorted by topic.HUB NL AI and automationThe service behind it: from individual skills to working automation in your business.SRV AI trainingIf your team wants to learn to set up and sustain this kind of thinking itself.SRV NLFrequently asked questions
What does the AI First Principles Thinker skill cost?
Nothing. The skill is free, falls under the MIT licence, and you do not have to create an account or leave an email address. You download a 6.9 KB zip containing a folder and a single file, SKILL.md, and that is the complete skill. There is no paid version and no sales email follows.
Is first principles thinking the same as reasoning from first principles?
Yes. First principles thinking and reasoning from first principles refer to the same way of reasoning: reducing a plan to what is demonstrably true and building it back up from there. The file also responds to thinking from the ground up, getting back to the core, assumption mapping, riskiest assumption, breaking down a cost structure, and blind spots.
Does this skill also work in Codex, Cursor or Gemini CLI?
Yes. SKILL.md has been an open standard since December 2025, so the same file also works in Codex, Cursor, Gemini CLI and other tools that follow the standard. In Codex you unzip it into .agents/skills/ in your project, or into ~/.agents/skills/ for all your projects; Codex has supported SKILL.md directly since the open standard of December 2025. Putting the contents of SKILL.md into your AGENTS.md still works too. The instructions are just readable text, so any AI assistant that accepts instruction files can handle it.
How does the skill score my assumptions for risk?
It surfaces eight to twelve assumptions, spread across the customer, the market, execution, technology, money and regulation, and writes every assumption as a claim that could turn out to be wrong. Every assumption then gets two separate figures from 1 to 5: the probability that it is wrong and the impact if it turns out to be wrong. The product of those two sets the order, and for every assumption the skill also records the evidence strength: proven, plausible or unproven. Merging probability and impact into a gut-feel score is on the list of things the skill never does.
What are primary risks and why a maximum of four?
The primary risks are the assumptions with the highest combined score, usually two to four of them. For each risk the skill delivers two things: the cheapest test you can run this week to find out whether the assumption holds, and an adjustment to the plan that shrinks the risk even if the assumption does hold. It refuses to flag more than four primary risks, for the reason that treating everything as important makes nothing important.
How does this skill differ from the 5 whys method?
The 5 whys method digs downward from a problem that has already occurred and looks for the cause underneath it. First principles thinking does not go to the cause but to the foundation: it strips a plan down to what is demonstrably true, scores the assumptions that remain for risk, and builds back up from there. The first is diagnosis after the fact, the second is groundwork beforehand. They do not conflict and are often used one after the other.
When is it better not to use first principles thinking?
The file is clear about this itself. Not for routine jobs where the approach is settled and the costs are low, because thinking from the ground up costs time, and you only earn that time back on decisions that are hard to reverse. And not if you just want a decision confirmed, because the skill confirms nothing out of politeness.
From score table to test
The temptation after a breakdown session is strong: tear up the whole plan straight away. Do not. The result is a short list of claims your plan rests on, sorted by how hard they could hit you. Take the top two and run the cheapest test listed under them. Fifteen conversations, a €200 ad set, a look at figures you already have: it is precisely the small tests that turn the score table into something you can act on. Only once a test like that is in does the evidence strength move from unproven to proven.
If one of your assumptions is about visibility, for instance the assumption that customers will find you anyway, that one is easy to test. The free SEO scan shows within seconds where your site stands, and that is evidence instead of a feeling. And if you want to talk through what else AI can do for your business, from individual skills to full automation, we simply do that in a conversation.