The earnest builders
Hunter, thanks for joining us. An AI researcher has resigned, and you’ve asked to make a statement.
Yes. I understand it’s an exceptionally serious situation. Someone has left a job.
Jacob Coxon has left Anthropic. His warning is rather bigger than the resignation.
That’s what interests me, Jon. People die in a disaster and the editor needs a death toll before deciding how much space it deserves. An AI man hands in his notice and we’re immediately discussing the survival of the species. Extraordinary return on a resignation letter.
If the people building the technology think it could kill everyone, that deserves attention.
Which people?
Let’s take his wording: “The people building AI earnestly believe that it could kill us all by the end of the decade.”
Right. Here’s how I read that. “The people building AI earnestly”—those people—“believe that it could kill us all.”
The earnest builders.
Exactly. Not necessarily everyone building AI. The people doing it earnestly. Whoever they are.
I initially read “earnestly believe”. The builders are unspecified; the belief is sincere.
And there we are. You’ve got earnest believers. I’ve got earnest builders. Same sentence.
So “earnestly” can reach backwards to “building” or forwards to “believe”.
Yes. In my reading, the subject is “the people building AI earnestly”. What do those people do? They believe. In yours, the subject is “the people building AI”, and what do they do? They earnestly believe.
One word doing two possible jobs.
Very efficient. Perhaps it’s already replaced someone.
Neither reading tells us exactly who these people are.
That’s the problem. Everyone building AI? The earnest builders of AI? His colleagues? People he had lunch with? We’re being asked to take the belief seriously because of who holds it, and we haven’t established who fucking holds it.
He’s invoking the authority of the builders.
Yes, but we need to know which builders before we can decide what their authority amounts to. “The people building AI” is a very large room to speak for.
He also says, “This is not a marketing stunt.”
Very reassuring. I hadn’t asked yet. Thoughtful of him to get ahead of it.
You hear a pre-emptive defence.
Of course. Anthropic would never derive any promotional benefit from the suggestion that its technology is so powerful it might destroy the species. Nobody could possibly hear that and think, “Fuck me, they must be building something impressive.”
The concern could be sincere and still serve that purpose.
Absolutely. I’m prepared to grant everyone their sincerity. It’s the most abundantly supplied component of the argument.
Old HaHu
And someone inside Anthropic did confirm that he holds the belief. Evan Hubinger.
Close personal friend.
Really?
Harbinger Hubinger, we used to call him. Old HaHu. HarbiHubi, if we’d had a few.
HarbiHubi? And he answered to that?
What do you think, Jon? What can I say? He’s a glass half empty kind of guy.
You can’t leave me with “Harbinger Hubinger” and no explanation.
Never thought he’d get where he has, to be totally honest, since that first time he declared the end of the world after listening to “Zombie” by the Cranberries. Though admittedly we were all on acid at the time. Ha! College, Jon, what are you gonna do?
Most people would have just turned the music down.
Always had that fatalist streak. As I said, good to see some people never change, even after all these years.
The forecast survived. The explanation’s had an upgrade.
At least back then he had the courtesy to blame the Cranberries instead of recursive self-improvement.
And now there’s a percentage.
That’s professional development, Jon.
For the record, this college history is your account.
And I’m expressing it earnestly.
“That’s professional development, Jon.”
Hunter Fitzpatrick
Linguistic judo
Let’s return to what we can read. Hubinger writes: “Jacob is correct here—we really do earnestly believe AI could kill all humans!”
There you are. Earnestness firmly attached to belief. “We really do earnestly believe.” No chance of it wandering back towards the building now.
He removes the ambiguity you found in Coxon.
Exactly. Coxon gave me earnest builders; Hubinger supplies earnest believers. The Institute takes a close interest in what happens to a statement during transmission.
Coxon’s talking broadly about the people building AI. He doesn’t say only Anthropic.
Exactly. Then Evan steps forward and says, “We really do earnestly believe.” And suddenly the room has an Anthropic sign over the door.
The broad group gets a particular company’s face.
Yes. Now go back to my first reading: the people building AI earnestly. Who are you picturing?
Anthropic.
There you fucking go. Nobody said they were the only earnest builders. But they’ve volunteered to embody the earnestness.
So the reply can colour how you remember the original sentence.
That’s the linguistic judo, Jon. You start with a warning about an industry and come away with an impression of one company’s character.
Even though the thing they’re sincerely telling you is that they might kill everyone.
And yet somehow I’m learning how responsible they are. Extraordinary communications work.
Then he says, “I personally think it is >10% within the next decade.”
So we’ve moved from what an unspecified group believes to one person’s numerical estimate. We’ve established that Evan believes it. How do we establish the probability?
He’s the alignment-science lead.
That establishes where he works.
And relevant expertise.
Certainly. A reason to ask him questions. Does it also answer them?
What would you ask?
Well, he goes on to say they don’t have a plan to solve alignment for superintelligence and aren’t clearly on track to get one. He’s the alignment-science lead, Jon. He leads it. What does the leading consist of under those circumstances?
Research into a problem they haven’t solved.
Fine. Then tell us how the estimate comes out of that research. What assumptions? What chain of events? What would make him revise it? We’ve got “I personally think” attached to the possible death of everyone. I’d like to see the working.
Starting with what he means by superintelligence?
Please. Apparently it has definitions. Plural. Which fucking one has a greater than ten per cent chance of killing me?
You want a definition stable enough to test the claim against.
Exactly. What system are we discussing? What can it do? How do we establish that it can do it? “Superintelligence” can’t do the work of an entire explanation.
He does say Anthropic is trying its best.
Oh, well. Put that in the headline.
A research culture
You seem particularly taken with his choice of words.
What I really love about Anthropic is how they combine em dashes with claims of being earnest whilst making unsourced claims about human extinction followed by exclamation marks. It really says something about their character as a company.
The exclamation mark has stayed with you.
He’s announcing it as though we’ve won something. I don’t need him to sound more enthusiastic, Jon. I need him to finish the thought. You believe this. You lead the relevant science. You don’t have a plan. So tomorrow morning, what do you do?
And what would count as a satisfactory source?
Something I can examine beyond his earnest personal assurance.
Suppose the source is supplied in a reply—and it takes you to another X post.
I fucking love that, Jon. That’s precisely the sort of research culture I’m looking for. At the Institute we’ve spent months building separate websites so we can cite each other. These people are doing it in the replies.
A considerable saving on infrastructure.
It’s really important for responsible AI safety companies to show leadership by maintaining a presence and publishing on platforms owned by trillionaires and dominated by the far right. It’s the most responsible information dissemination strategy I’ve seen since my time in sub-Saharan Africa working for Nestlé.
What was your role there?
The past is the past, Jon, and I didn’t sign that NDA for free.
Let’s stay with research practices, then. Perhaps the Institute should investigate microdosing.
Microdosing! They’ve turned taking drugs into another fucking productivity target. You used to take something and lose an afternoon. Now you’re expected to come back with a better spreadsheet.
A measurable return on the experience.
Even your altered state of consciousness has to justify its salary.
What’s the Institute’s policy?
Strict macrodosing.
Scheduled, or as required?
Whenever the peer review is looking particularly threatening.
Does it change the findings?
It changes our relationship with them. Secondary contextualisation.
“The past is the past, Jon, and I didn’t sign that NDA for free.”
Hunter Fitzpatrick
The end of the humanities
That’s an existing Institute procedure?
Yes. Findings that don’t fit the framework receive further interpretation before entering the evidential record.
And you’ve been applying it to these extinction forecasts?
We’ve developed an extension. Informally, its working title is the Chinese Whispers approach.
What does the extension add?
Another round.
Take me through it.
We begin with the people building AI and their unsettled relationship with earnestness. Then Evan confirms his own belief and supplies a percentage. In the reporting, killing all humans becomes the end of humanity. We apply further contextualisation and arrive at the end of the humanities.
Humanity to humanities.
A small textual adjustment with considerable scope for original research.
The first paraphrase can preserve the meaning. Yours changes it.
That’s our contribution, Jon. We can’t keep expecting Anthropic to do everything.
And the resulting prediction?
A 10% baseline probability of the end of the humanities by the end of 11 July 2029.
You’ve kept the percentage.
We felt it was important to preserve the scientific component.
Reduced it slightly. His was greater than 10%.
Ours is a conservative baseline.
You’ve also brought the deadline forward.
Coxon said the end of the decade. Hubinger said within the next decade. We’ve resolved that uncertainty by supplying a precise date.
Why 11 July 2029?
I felt “sometime” would undermine the rigour.
What would the end of the humanities look like?
People losing the ability to interpret language. To notice that a word can attach in two places. To ask who a statement refers to. To distinguish evidence that someone believes something from evidence that it will happen.
There could still be humans.
Oh, certainly. The rich will continue to dwell in their underground sugar caves. I’m concerned about what else survives. Humanism. Values. The capacity to be good. To understand that other people exist even when they haven’t accepted your terms and conditions.
Your protocol exploits the very failure you’re forecasting.
Which is why we’re taking it so seriously.
You’re demonstrating how authority survives changes in meaning.
Look how much has changed in this conversation. The people, the proposition, the deadline. But we still have a percentage and a director willing to stand beside it.
You.
And you’re publishing it.
As an interview.
We’ll cite it as external coverage.
So the Institute’s forecast acquires a NEPRAVDA reference.
And NEPRAVDA acquires an expert forecast. We give back to those we owe, Jon.
Does that make this a marketing stunt?
I would have thought my earnestness had settled that.
“We felt it was important to preserve the scientific component.”
Hunter Fitzpatrick
A decision of principle
Let’s return to the resignation. Whatever you think of the rhetoric, Coxon has acted on his concerns.
I mean, really, Jon. This guy, this so-called AI safety “expert”—what a genuine misanthrope. I’m telling you right now: he’s going to regret this move. Who on earth would leave a job in the AI industry? Does he understand what an opportunity he’s throwing away?
You’ve spent much of this interview questioning the company’s public claims. Now you’re appalled that he left.
He had a job at Anthropic, Jon. He was already inside the building. Whatever his concerns were, he could have continued having them on a very good salary.
Like your old friend HarbiHubi.
Exactly. Still there. Still concerned. You have to respect the consistency.
You think Coxon will miss it?
As I said: misanthrope. And he’s going to miss Anthropic for sure.
Would you have stayed?
I would have appreciated the opportunity.
That sounds suspiciously like something from a covering letter.
I take a close interest in their research culture. I have directly relevant experience in secondary contextualisation. I’m comfortable publishing findings whose supporting material is available somewhere else. And I know the Harbinger.
You’re unusually well prepared for a man who came here to discuss the humanities.
You have to think about your own future as well as everyone else’s.
Before we finish, what should readers actually do with these warnings?
Read carefully. Follow the references. Ask whether the source supports the claim. Notice when someone moves from what other people believe to what you’re supposed to believe. And consider whether the person speaking might want something.
I was coming to that. What do you want, Hunter?
I’ve thought very seriously about my position. You can’t keep talking about values without eventually making a decision that affects your own life.
And that’s why today I’m announcing my resignation.
From the Shit Science Institute?
Yes.
Because of your concerns about the future of the humanities?
Because apparently there’s an opening at Anthropic.
Correction · 12 September 2026: An earlier version of this interview omitted a link to the SSICED paper accompanying Hunter Fitzpatrick’s account. The source material, The Death of the Humanities, is now linked here. The paper also cites this interview.
The public record
Original posts discussed in this interview.
Jacob Coxon’s resignation statement ↗Evan Hubinger’s response ↗