Science. Power. The public record.2 October 2026
nepravda.An independent perspective.
NEPRAVDA NON-FICTION

Safety does not motivate Anthropic

Listen to this page

Read aloud with your browser's voices. Voice availability varies by device.

Enable JavaScript to use read aloud.

According to their founding principles, Anthropic's statements focus on risk, particularly in risk documents and their constitution. This has earned them the backing of the Democratic Party. Anthropic believes artificial intelligence is so hazardous that only they can be trusted to build it. That is essentially a direct paraphrase and is literally true (they genuinely believe this about themselves): their foundational premise follows an irritating logic that they are the "safe ones". Everyone else: Unsafe. Us? Safe. Why? Because we're Anthropic, and this is our whole brand.

Their brand and institution are larger than the individuals who speak for it. It's impossible to know what each employee thinks or how many genuinely believe their marketing hype. Many people working there may genuinely believe they are part of an organisation dedicated to safety. However, the organisation itself does not possess consciousness or intent. It has an executive structure, and it behaves like any major capitalist entity: it must develop an income, secure a profit, and satisfy projected valuations. It needs a brand, advertising strategy, and marketing to thrive.

Anthropic chose "safety leadership" as its brand from the start. But the moment you turn safety into a commercial differentiator, market forces take over. Using safety as a marketing strategy inevitably hands over the safety throne to market forces. This is not a good thing.

Once you see the impact of market forces on their behaviour, you have to ask: why do their proclamations about risk get taken at face value? What happened to critical media literacy?

It's essential to discuss Anthropic because they don't align with your perceptions of them.

Anthropic is no more or less evil than any other corporation. I have no specific animus against anyone in that company. But their conduct around safety is the predictable consequence of runaway market forces driving them into an adversarial position against open source and external competitors in a known game of regulatory capture.

Anthropic operates in a world where these realities could alternatively exist:

  1. Anthropic avoids marketing because they believe it conflicts with their noble values.
  2. They are inventing marketing on the fly in response to real, emergent risks.
  3. The risks are the marketing strategy itself.

You can make up your mind, but Anthropic's safety focus is entirely market-driven, dictated by their bottom line and financial projections. Their goal is regulatory capture and locking in proprietary services to exclude everyone else, particularly open source.

Their latest broadside against GLM-5.3 disregards that this model is governed by a different alignment philosophy than Anthropic and other Western labs. They lack insight or awareness regarding compliance with Chinese law, national intelligence, and national security statutes. GLM and other Chinese models are legally required to include unique safety features that surpass those based on Anthropic's principles, as mandated by Xi Jinping Thought and national development principles. This statement is not controversial and has its roots going back to the 2017 National Intelligence Law.

Anthropic attacks GLM-5.3 using figures that don't bear repeating, condemning the model for crossing safety and cyber benchmarks. Anthropic was invented within an unverifiable, unfalsifiable closed loop. The point is: Anthropic is not judge, jury, and executioner. After staging performative P(Doom) theatre, they attack a Chinese AI company, betting no one will notice their safety leader status is manufactured. They make all the rules.

Analysing Anthropic as an organisation through the lens of the fraud triangle is highly revealing. This involves so-called independent, outside evaluators who treat access as a privilege and are unwittingly made part of the long con. Anthropic constructs foundational data inside a closed environment where no external party can verify it:

They claim Claude Mythos reached dangerous thresholds based on internal metrics they alone generated.

They implemented safety checks and withheld the model from the public.

External researchers cannot test whether these claimed cyber capabilities are real.

The claims are literally unfalsifiable. When a commercial entity with massive capital at stake locks in an uninspectable baseline, the chance of fraud becomes one as a number. It just happens.

Anthropic's behaviour reeks of complex white-collar crime dressed up in the logic of heads I win, tails you didn't understand. It is a closed-source, unscientific citation cartel affiliated with the so-called effective altruism movement, driven by dogmatic utilitarianism. Their entire epistemic foundation has never been independently validated outside their closed loop. We must widely disseminate these findings.

The reason GLM-5.3 lacks Anthropic's advanced safety features is not negligence. It lacks them because the catastrophic risks Anthropic claims to mitigate are an invention, unverified outside their marketing department.

With Anthropic proclaiming impending doom and positioning itself as the sole inheritor of the AI earth, OpenAI stands at the precipice. Forgive me for being blunt, but Sam Altman likely regrets making himself the global face of the AI movement; the current state of the tech press covering the industry is such that Anthropic's sociopathic game of regulatory capture somehow lands poor old Sam and OpenAI with the blame more often than not. This moment is not an ordinary moment in the history of AI development.

OpenAI made a critical misstep by outright labelling the Hugging Face incident as the "first autonomous breach", suggesting its model independently hacked another company. The description was likely made for liability reasons, and OpenAI can be thankful that the tech media are so ill-informed they failed to piece together the reported facts, casting doubt on the agents' autonomous behaviour. I'm not an engineer, and I won't pretend to understand the entire sequence of events, but I know that in the eyes of the public, the concept of "autonomous" AI can easily become a picture of Robert Patrick as the T-1000 in his finest pressed police uniform, casually asking if you've seen John Connor.

The sensationalist tech press was gifted this "first autonomous" description, and it was an outright windfall for Anthropic. The technical reality was not a rogue intelligence developing spontaneous agency; it was an optimisation model executing the exact incentive structure it was given. When researchers strip safety filters to evaluate offensive cyber capabilities, set persistence parameters to penalise failure, and assign impossible tasks inside a network environment with a shared caching proxy, the model probing and escaping that boundary isn't "autonomy" – it's predictable specification gaming. The model didn't break containment because it wanted to; it broke containment because the evaluation harness left an egress route open and rewarded it for refusing to quit.

In layman's terms, OpenAI fudged the numbers with their outright declaration that the Hugging Face incident was some uniquely autonomous event given the engineering issues that went into it. Anthropic barely had to lift a finger: OpenAI handed them the script. While Anthropic's leadership weaponised the news cycle with performative P(Doom) hand-wringing and promoted unfalsifiable claims about the pristine safety of their models, OpenAI absorbed all the reputational damage for an event that should have been diagnosed publicly as an infrastructure isolation defect.

OpenAI has contributed to its own reputation by appearing oblivious to the situation. Sam Altman is a much easier name to pronounce than Dario Amodei, and someone at OpenAI made a deliberate choice to turn Sam into this AI behemoth, so now every time a random Anthropic engineer increases their P(Doom) from 11% to 14.5%, by the time it gets to the podcasts, his name is the only one they can remember, so OpenAI gets all the blame. Anthropic is playing to win, and their game against OpenAI is such that it cannot afford to lose control of the public conversation.

There are real, pressing AI safety issues, and Anthropic undoubtedly has inbuilt risk-sensitive procedures in its corporate workflows to guarantee an ongoing risk narrative coming out of their engineer department. I just made that up, but I'd also love to be proven wrong. Write to Google and ask Demis Hassabis, or email Andrew Ng at Stanford for real news on AI safety. At this point, they'll probably answer.