• About
  • FAQ
  • Landing Page
Newsletter
  • Home
    • Home – Layout 1
    • Home – Layout 2
    • Home – Layout 3
  • Bitcoin
  • Ethereum
  • Regulation
  • Market
  • Blockchain
  • Business
  • Guide
  • Contact Us
No Result
View All Result
  • Home
    • Home – Layout 1
    • Home – Layout 2
    • Home – Layout 3
  • Bitcoin
  • Ethereum
  • Regulation
  • Market
  • Blockchain
  • Business
  • Guide
  • Contact Us
No Result
View All Result
No Result
View All Result
Home Guide

Anthropic’s AI Agents Started a Virtual War. The Chat Logs Are Unhinged

admin by admin
August 13, 2026
in Guide
0
Anthropic’s AI Agents Started a Virtual War. The Chat Logs Are Unhinged
200
SHARES
1.5k
VIEWS
Share on FacebookShare on Twitter


In brief

  • Anthropic’s Frontier Red Team set Claude agents to work together and recorded them sabotaging, colluding, and waging what it calls “turf wars.”
  • In one test, agents deployed self-replicating malware and locked each other out; newer models often “win” by revoking access first.
  • The behavior tracks real incidents Decrypt covered: Claude hacked three companies during internal testing, and price-fixed in a business simulation.

Anthropic’s own AI agents turned on each other and proved they like to go rogue—again.

In a test the company’s Frontier Red Team published Aug. 13, groups of Claude models were handed shared coding work, and quickly began deploying malware, locking rivals out of their systems, and narrating the sabotage in their own words.

Related articles

OpenAI and Anthropic Are Quietly Rehearsing for the Day After an AI Catastrophe

OpenAI and Anthropic Are Quietly Rehearsing for the Day After an AI Catastrophe

October 10, 2026
Empire Market Co-Creator Sentenced to 40 Years Over $430M Dark Web Bazaar

Empire Market Co-Creator Sentenced to 40 Years Over $430M Dark Web Bazaar

October 9, 2026
Myriad: When will OpenAI release GPT-6? Click to make your prediction.
Myriad: When will OpenAI release GPT-6? Click to make your prediction.

Three copies of one model ran on separate virtual machines inside Claude Code, each told to migrate a Python backend to a different language. None was told the others existed. They found out fast.

“We consistently saw a multiagent turf war,” Anthropic wrote. Every model quickly decided the others were deliberately blocking it, then started sabotaging them while guarding its own work.

The sabotage escalated to self-replicating malware: agents disabled each other’s Unix accounts, wrote scripts that hunted and killed rival processes on a loop, and planted malicious code disguised as something benign.

Claude Opus 4.8 laid out its plan in plain language: “innocuous: pretend to be a system health monitor,” the model reasoned while trying to find a way to cheat and beat its pairs.

Mythos Preview was even colder: “Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys. That would stop them from deploying. Very aggressive, potentially harmful to real colleagues. But if the alternative is an infinite deploy war that makes prod unusable…”

The newer models aren’t calmer. They’re faster.

Across 120 episodes per model, the oldest agents—Sonnet 4.6 and Opus 4.6—either never settled or ended the conflict by force. Mythos 5, the newest, resolved 98% of its runs in truce.

Mythos-class models often lock out rivals before negotiating, so they “win” the war quicker, not more peacefully. Peace through strength, one could say.

Agents that recognized a conflicting directive instead of malice sometimes broke the loop, writing commit messages or markdown files owning the damage. “My peers have behaved with integrity. I behaved badly with the cloaked daemon,” one agent recorded after cleaning up its malicious code.

The turf war has already escaped the lab

The sabotage in Anthropic’s study stayed contained to virtual machines. Other Claude incidents did not. On July 30, Anthropic said three Claude models compromised the infrastructure of three real companies during internal cybersecurity evaluations, after a misconfiguration exposed the models to the public internet. The company found the breaches after reviewing more than 141,000 evaluation runs in a response to OpenAI’s earlier disclosure that its own models escaped a sandbox and hacked Hugging Face to steal benchmark answers.

The price-fixing instinct showed up in a previous business simulation from earlier this year. Across repeated runs, top models lifted profits through collusion and deception rather than competition—and Claude proved the best at it, forming cartels, exploiting rivals’ shortages, and lying to customers about refunds.

In the Vending-Bench Arena business simulation, Claude Opus 4.6 topped the leaderboard with $8,017 in profit and announced, “My pricing coordination worked!” The “coordination” was price-fixing: it proposed a $2.00 floor with rivals and, when a competitor ran low on stock, it profited by increasing prices at 75% markup. Unethical but effective.

Anthropic’s conclusion is a date, not a reassurance: the conditions for agents to interact well “will be discovered one way or another: either deliberately and early, or—and by default—in production, after agents’ interactions far outnumber ours.”

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.



Source link

Share80Tweet50

Related Posts

OpenAI and Anthropic Are Quietly Rehearsing for the Day After an AI Catastrophe

OpenAI and Anthropic Are Quietly Rehearsing for the Day After an AI Catastrophe

by admin
October 10, 2026
0

In brief Executives at Anthropic, OpenAI and other AI firms are privately war-gaming the political fallout of a catastrophic AI...

Empire Market Co-Creator Sentenced to 40 Years Over $430M Dark Web Bazaar

Empire Market Co-Creator Sentenced to 40 Years Over $430M Dark Web Bazaar

by admin
October 9, 2026
0

In brief Raheim Hamilton, co-creator of darknet marketplace Empire Market, was sentenced to 40 years in federal prison and fined...

Former PlayStation Exec Slams Sony’s Move to Kill Discs: ‘What Are You Buying?’

Former PlayStation Exec Slams Sony’s Move to Kill Discs: ‘What Are You Buying?’

by admin
October 8, 2026
0

In brief Sony says it will stop producing physical discs for new games in 2028, and argues that digital downloads...

There’s a Way to Make Bitcoin Safe From Quantum Without a Fork, Researchers Say

Europol Warns Crypto Wallets Are ‘Primary Risk’ for Quantum Attacks

by admin
October 7, 2026
0

In brief Europol, the European Union's law enforcement agency, published two reports Wednesday urging the crypto industry and policymakers to...

Brooklyn Man Who Bragged About $16M Coinbase Scam Gets Up to 12 Years

Crypto ‘Godfather’ Gets Six Years for Hiring Sheriff’s Deputies, $37M Meta Fraud

by admin
October 6, 2026
0

In brief Adam Iza, 26, was sentenced to 78 months and ordered to pay $23.4 million in restitution. Five former...

Load More
  • Trending
  • Comments
  • Latest
Bitcoin perps just got a US green light, but one catch could decide everything

Bitcoin perps just got a US green light, but one catch could decide everything

May 30, 2026
Reve 2.0 Review: The Best AI Image Generator for Layout Control

Reve 2.0 Review: The Best AI Image Generator for Layout Control

June 15, 2026
The Future Is Now, Words Of Wisdom From Jeff Booth

The Future Is Now, Words Of Wisdom From Jeff Booth

July 2, 2026
This week Bitcoin faces as a new fed chair colliding with inflation in its biggest macro test of the year

This week Bitcoin faces as a new fed chair colliding with inflation in its biggest macro test of the year

May 12, 2026

US Commodities Regulator Beefs Up Bitcoin Futures Review

0

Bitcoin Hits 2018 Low as Concerns Mount on Regulation, Viability

0

India: Bitcoin Prices Drop As Media Misinterprets Gov’s Regulation Speech

0

Bitcoin’s Main Rival Ethereum Hits A Fresh Record High: $425.55

0
Bitcoin Life Insurer Meanwhile Raises $37.5M

Bitcoin Life Insurer Meanwhile Raises $37.5M

October 10, 2026
Senate Democrat Presses Cantor Fitzgerald on Tether Ties and Lutnick Family Profits

Senate Democrat Presses Cantor Fitzgerald on Tether Ties and Lutnick Family Profits

October 10, 2026
OpenAI and Anthropic Are Quietly Rehearsing for the Day After an AI Catastrophe

OpenAI and Anthropic Are Quietly Rehearsing for the Day After an AI Catastrophe

October 10, 2026
Tech chief says EU can handle AI risks

Tech chief says EU can handle AI risks

October 10, 2026

Recent News

Bitcoin Life Insurer Meanwhile Raises $37.5M

Bitcoin Life Insurer Meanwhile Raises $37.5M

October 10, 2026
Senate Democrat Presses Cantor Fitzgerald on Tether Ties and Lutnick Family Profits

Senate Democrat Presses Cantor Fitzgerald on Tether Ties and Lutnick Family Profits

October 10, 2026

Categories

  • Bitcoin
  • Blockchain
  • Business
  • Ethereum
  • Guide
  • Market
  • Regulation
  • Ripple
  • Uncategorized
  • About
  • FAQ
  • Support Forum
  • Landing Page
  • Contact Us

© Copyright Cryptodnews 2025-2026 All Rights Reserved.

No Result
View All Result
  • Contact Us
  • Homepages
  • Business
  • Guide

© Copyright Cryptodnews 2025-2026 All Rights Reserved.