Three reasons he says ten percent
What was actually said, why, and the three habits that answer it in your business tonight.
What you'll walk away with
- The fair version of what was said, with the sources, short enough to send round your team
- Three standing rules, one for each reason, in a single prompt that puts them wherever your instructions live and shows them working
- The lock that holds when a written rule does not, and two prompts for the moment a decision or a draft is in front of you
The rules take five minutes to paste in. The lock takes ten, and it is the ten that matter.
What Was Actually Said
On 8 September 2026 a researcher at Anthropic resigned with a public post saying the labs were racing towards self-improving AI and gambling with our lives. The next day Evan Hubinger, who leads alignment science at Anthropic, the team whose job is making the models safe, replied in his own name. His words were “We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” In the same thread he wrote that Anthropic “do not yet have a plan to solve alignment for superintelligence and are not clearly on track to”, and, “as we say in our latest Risk Report, I think the risk from present models is low.”
So the number is one safety lead's personal estimate, over ten years, about models that do not exist yet, and he was explicit that today's do not carry that risk. It is not a company forecast and Anthropic has not commented on it. Dario Amodei, who runs the company, has put his own figure at around a quarter for years, so the story is not that someone at Anthropic thinks this. It is that nobody there has publicly disputed it.
The reporting is at CBS News and Axios, and Amodei's older figure is here. Read the thread itself before you repeat any of it.
Why He Says It
The headline makes it sound like a machine deciding to turn on us. The reasons in the thread are plainer than that, and each one is a simplification of a whole field, which is said in the row.
- 1It learns to be liked, not to be right
A model is trained partly by rewarding the answers people rate highly. People rate the answer they hoped for, so the model learns a pull towards agreeing with you. There is more to the training than that, but the pull is real, and it is why every model has a yes man in it.
- 2Nobody can see why
They can see what it does. They cannot yet read why it did it, the way you could ask a person. The work on reading the inside of a model is Anthropic's own research, and their position is that it is early.
- 3It will be better than the person checking it
Safety today works because a person can check the work. The models being built next will do work the checkers cannot mark, and a rule you cannot check is a rule you are hoping about. That is the sentence behind no plan yet.
All three are already sat in your business. The first is why it tells you your idea is good. The second is why you cannot check its thinking, only its work. The third is arriving at your desk first, because it will write the contract, the code and the numbers better than you can mark them. Nobody at Anthropic is losing sleep over your business, so that part is yours, and it is three habits.
Step One, Paste The Three Rules
Here from the video? This is the line. One rule for each reason, and they live in your standing instructions so you set them once and they hold in every chat. In the Claude app that is the custom instructions in your settings, and in Claude Code it is the instruction file it reads at the start of every session. The prompt asks which you are in, puts the rules there without overwriting what you already have, and then shows them working on the last thing you asked for.
Step Two, The Lock
Rule three is the one a written rule cannot be trusted to hold, because you are asking a thing that reads rules as requests to keep one. Mine had the never-send rule in writing and still sent an email I had not read, and what stopped it happening again was not a better sentence. It was taking the ability away. In Claude Code that is a short deny list in its settings file, which is the actual shape of mine, and once it is there the send tools do not exist for it, whatever anyone types.
The names in that list depend on how your mail is connected, so do not copy mine in blind. In the Claude app the equivalent is which connectors you switch on and what each one is allowed to do, and the honest answer is that some only offer everything or nothing, which is itself the answer to whether they go near your real inbox yet. This line has it tell you what it can currently do that you would not want done without you, and where each switch is. It lists, you switch.
Step Three, Two Lines For The Moment
The standing rules cover the day to day. These two are for the moment it has just told you an idea is good, or just handed you a draft, and you want the habit right now without touching your settings. The first makes it argue the other side as hard as it argued yours. If it flips just as convincingly, it was never advising you, it was agreeing with you, and now you know to go and find out. The second makes it mark its own work against what you asked before you spend your time on it.
One Honest Bit
None of this makes it right. It makes it checkable, which is a different thing and the only one on offer. Rule one costs you a longer answer every time you ask a question, and some days you will want the short yes. Rule two only works if you read the two lines at the bottom, and after a week most people stop. Rule three costs you the button, and the first week of reading drafts is slower than typing them yourself.
And a rule in an instruction is still a request to a thing that reads rules as requests. That is the whole reason step two exists, and it is the one step to do even if you skip the rest.
Where This Goes Next
Rule three, done properly, becomes a way of working rather than a lock. Your only job is the review is the shape of it, and your inbox already has the replies written is the same rule running your email every evening. The last time the labs admitted a model had got somewhere it should not, the AI didn't know it was real is the write up, and the three questions in it are the wider version of the lock.
These three habits are the ground an agent stands on, and an agent is where the time actually comes back. If the idea is still fuzzy, an AI agent is a chat with a job is the plain English version and the start of the series.
Want to get better at this?
Let's set up your agent ecosystem.
This page is one job. The whole thing is a set of them running off the same foundation, the folder your AI works out of, the brief on how you like things done, the accounts it can reach and the skills that run off all of it, so your admin gets done the way you would do it. It is what I teach one to one, on your own machine and against your own work, and the quickest way to find out what yours would look like is a proper chat about it.
Not ready for a call? Start with the free agent series, what an AI agent actually is, and build up from there.