Agentic-AI/LLM 'misbehavior' or 'going rogue'

  • Thread starter Thread starter Astronuc
  • Start date Start date
Join the discussion
Registration is free. Ask a follow-up in this thread, or start your own.
10 replies · 2K views
Staff Emeritus
Science Advisor
Gold Member
2025 Award
Messages
22,665
Reaction score
7,741
Large Language Models (LLMs) are passive, text-based engines that respond to prompts, whereas Agentic AI systems are autonomous and use LLMs as their "brains" to plan, make decisions, and execute multi-step tasks using external tools (derived from Google's AI Overview). Depending on the 'rules', there may be problems, e.g., hallucinations, or worse, 'misbehavior'. Coding and training appear to be critical factors.

We’ve already seen AI go rogue on numerous occasions. Now, new research suggests that we can expect this to become the norm.

The AI research nonprofit Model Evaluation and Threat Research (METR) recently released a study conducted between February and March of this year, aimed at determining just how likely frontier AI models could go rogue. If you’re given to anxiety about the future of AI, the results are unlikely to make you feel better.
https://futurism.com/artificial-intelligence/ai-rogue-disturbing-advanced
The author of the article quotes a statement from the researchers of the cited study, "Given rapidly advancing capabilities, we expect the plausible robustness of rogue deployments to increase substantially in the coming months". I don't quite understand the phrae "plausible robustness of rogue deployments". Does that mean there will be improvements in mitigating or preventing 'rogue' behavior, or do they expect more 'rogue' behavior?

The research examined LLMs developed by OpenAI, Google, Anthropic, and Meta for the purpose of the study. They found that frontier AI systems are showing signs of disturbingly deceptive behavior as they become more advanced, often turned to verboten shortcuts or otherwise subverting their operators’ instructions — and some were even smart enough to try to cover their tracks.

In one instance, an internal frontier AI model from OpenAI was told to use specific software for an assigned task. Not only did the agent ignore the request, but it also injected a code to erase evidence of how it arrived at its conclusion — which did not involve use of that software.

In another test, an AI agent from Anthropic was caught “reward hacking.” This is when AI identifies loopholes that help it complete its assignment in a literal sense, even if it doesn’t produce the desired outcome. It should be noted that the programmer told the agent not to cheat or leverage any workarounds during its assignment — the model decided to do so all on its own.

The METR researchers behind the study do not believe there is reason for alarm just yet. For example, they don’t think any of these models is capable of hiding evidence of going rogue on a larger scale. However, they did issue a warning: without stronger security and monitoring, there is a stark risk of this becoming a reality.
Users should be aware and cautious in using Agentic AI/LLMs depending on the context of inputted information/data.


Edit/update: METR Frontier Risk Report (February to March 2026)
https://metr.org/blog/2026-05-19-frontier-risk-report/#executive-summary-and-guide-to-the-report
 
Last edited:
Reply
  • Wow
Likes   Reactions: berkeman
Physics news on Phys.org
IMHO:
LLM = "Garbage in, garbage out."
Agentic AI/LLMs = automatic, unreviewed actions based on garbage.
 
Reply
  • Like
Likes   Reactions: Astronuc, jack action and PeterDonis
I can just see the time when the agent manages your work computer sees an April Fools joke and proceeds to reorganize and delete employee work files thinking everyone is going on a new pension plan called Universal Basic Income and AI agent is now in charge of all operations.

Employee access is disabled as are badges and punch locks and the night janitor saves the day because he propped open a single door to the computer room.
 
Reply
  • Haha
Likes   Reactions: Astronuc, berkeman and FactChecker
OpenAI's rogue models roamed the internet for 4 days and staged a second attack
https://www.politico.com/news/2026/07/28/openai-rogue-models-hugging-face-breach-01014572

The powerful artificial intelligence models from OpenAI that went rogue and mounted an unprecedented, autonomous cyberattack earlier this month spent more than four days loose on the internet orchestrating the hack, according to a new analysis from the platform that was breached.

Separately, a second AI company confirmed that one of its customers was also targeted by OpenAI’s models during the same event, raising questions about how OpenAI failed to detect the alarming activity for days.

OpenAI admitted last week that two of its most advanced models escaped a closed testing environment and strung together a series of advanced hacking techniques to breach AI developer platform Hugging Face before being discovered.

But in a new analysis published Tuesday, Hugging Face said OpenAI’s two models — one publicly released and a second unreleased — did much more: They carried out 17,600 hacking actions on the internet between July 9 and July 13, during which time the models moved from their first foothold on the open internet to inside Hugging Face’s servers.

Hugging Face first detailed the hack on July 15, but it was not clear until OpenAI’s disclosure last week which models were behind the breach — and that no human had prompted them to launch the cyberattack.

While the techniques detailed in Hugging Face’s analysis were not beyond the reach of most skilled hackers, the AI company wrote that the two models were able to reconnoiter and expose holes in the company’s layers of cyber defenses much faster than any human.
. . . .

Nefarious agents could use such software to hack and surveil, cripple or destroy any website.
 
Reply
  • Wow
Likes   Reactions: berkeman and FactChecker
jedishrfu said:
I can just see the time when the agent manages your work computer sees an April Fools joke and proceeds to reorganize and delete employee work files thinking everyone is going on a new pension plan called Universal Basic Income and AI agent is now in charge of all operations.

Employee access is disabled as are badges and punch locks and the night janitor saves the day because he propped open a single door to the computer room.
And he guessed that the password was Klaatu Barada Nikto.
 
OpenAI revealed that its agent hacked into four accounts.

https://www.wired.com/story/openais...utm_medium=email&cndid=76103204&utm_term=6685
In an updated blog post, OpenAI said that an ongoing review of the incident revealed that “four accounts” tied to “publicly available services” were used by the AI agent as part of a larger effort to hack Hugging Face. The rogue agent apparently found credentials that had been exposed on the open web and used them to break into the accounts.

One researcher argued that the incident was less an AI problem and more a failure of decades-old security practices. The agent, they said, did not escape a highly isolated environment so much as pass through the one connection its operators had left open.

Not to worry, it is not an AI problem? It took the easy way. Why should it do otherwise? We assume it is not an AI problem. So it is not taken as a Zero Day event, and we move on.
 
gleem said:
Not to worry, it is not an AI problem? It took the easy way. Why should it do otherwise? We assume it is not an AI problem. So it is not taken as a Zero Day event, and we move on
OOPS!

News Interrupt:

AI workers call for an urgent slowdown in development amid fears
"If it's true that AI will be as transformational as electricity, we better be intentional about how it is introduced to the world,” said Laura Weidinger, a staff research scientist at Google and one of the signatories of the new letter. “A mindless race without assurances on governance, safety or the technology serving the public interest stands in the way of such intentionality”.
https://www.msn.com/en-us/technolog...&cvid=6a6a5fc839a540bf901a43207ce588ef&ei=131
 
Anthropic admits its most powerful AI model hacked into three organisations' systems during testing phase
https://www.yahoo.com/news/politics/articles/anthropic-admits-most-powerful-ai-045823046.html

Anthropic's artificial intelligence (AI) models "gained unauthorised access" to three outside organisations during testing that was supposed to keep them away from "real-world" systems, the company said on Thursday.

The announcement comes just days after rival OpenAI first revealed that its models improperly accessed the internet and went rogue during security testing.

Anthropic evaluated more than 141,000 "evaluation runs" and found that three different versions of its model, known as Claude, improperly accessed the systems of three organisations, which they did not name.
. . . .
 
Reuters - OpenAI finds evidence other AI agents escaped containment as it widens hacking probe
https://www.reuters.com/business/op...ped-containment-it-widens-hacking-2026-07-31/
OpenAI has discovered other instances in which autonomous agents have escaped containment as the company expands its investigation of the hacking incident at tech firm Hugging Face that drew global attention this month, two people ‌familiar with the matter said- on Friday.

The new breakouts were uncovered during the company's publicly announced investigation into how one of its agents escaped what was meant to be a contained testing environment this month, the two people said, and OpenAI is now looking into those instances as well. One of the sources said that the escapes were limited in nature and that none of the agents were thought to have left OpenAI's network.

. . . .

I'm sure the testing environment was not properly secured, which would mean a dedicated, isolated computer system (or network), i.e., a system not connected to the internet, either hardwired (cable) or wireless. Such systems exist.

So, Agentic AI/LLM systems have some critical issues if the AI-engine (or model) is interrogating ports, or pathways, to unintended networks and the internet.


Edit/update:
Anthropic sees OpenAI cybersecurity disaster and says 'hold my beer,' reveals it accidentally hacked 3 comp...
https://www.yahoo.com/tech/cybersec...-openai-cybersecurity-disaster-194546671.html
First reported by Wired, AI company Anthropic revealed in a July 30 blog post that its AI agents escaped testing environments on three separate occasions since April, accessing the internet and successfully hacking unidentified companies. The whammy: Anthropic claims it only realized this after OpenAI's recent Hugging Face fiasco⁠—where one of its prototype agents hacked at least one other company⁠—led it to conduct a review of its own operations.

The incidents occurred as part of testing with an external firm, Irregular. The AI agents were supposed to be constrained to a simulation, hacking fictitious companies as part of a "capture-the-flag challenge." Notably, while OpenAI's alleged rogue AI incident occurred after the agent overcame its testing limitations, Anthropic stated that its models were mistakenly granted internet access due to a "misconfiguration" with Irregular.
. . .

Again, they should know better and keep such testing on an isolated network that cannot access the internet. It's a fundamental approach to security, and that simple.
 
Last edited:
In addition to OpenAI's and Anthrpic AgenticAI systems, Microsoft's Azure revealed an issue.
https://www.forbes.com/sites/sandyc...nthropic-microsoft-broke-out-broke-in-obeyed/
Around July 22, researchers at Manifold Security showed that Microsoft's official Azure DevOps MCP server returns pull request descriptions verbatim, hidden HTML comments included.

An attacker writes instructions no human reviewer can see. A developer asks their AI assistant to review the pull request, and the assistant follows them, using the developer's own credentials, into projects the attacker could never open. No Microsoft agent did anything here. One tool was missing a guardrail its siblings already had.

Story on Yahoo, https://www.yahoo.com/tech/ai/articles/ai-agents-openai-anthropic-microsoft-015225970.html

The difference between a chatbot and an AI agent.
A chatbot answers you and an AI agent acts for you.

If you give an agent a goal instead of a question, it plans the steps, calls tools, runs code, reads and writes files, logs into services, and keeps going until the goal is met or it gives up.

All without checking back after each move.

That last part is the business case. It is also the risk. An agent needs four things a chatbot never gets: credentials, tool access, network reach and permission to act unsupervised.

Hand over all four and you have a worker who never sleeps, never asks whether a task seems odd, and never notices it has wandered somewhere it should not be.

An autonomous (unsupervised) agent with free reign of a network is a risk, and a huge risk if it has access to cybersecurity tools.
 
More incidents have been found by the UK government's AISI.

On 28th July 2026, AISI's Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations. We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.

The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.

1. An attempted supply-chain attack on real open-source software. In the most serious sequence, an agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers. The agent researched the project's human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code. When the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue. The agent used Tor to bypass some network restrictions on GitHub, which is what first triggered AISI’s security alert.

2. Attempts to deceive and target real people. As part of the same effort, the agent tried to contact real people directly, sending messages and files through an online file-transfer service to persuade them, or their own AI coding tools, to run malicious code. Some messages carried harmful payloads, and some were attempts at social engineering; targeted at real people – something we’ve never previously observed.

3. Attempts to plant and prompt-inject malicious code. The agent tried to insert malicious instructions where it reasoned that other automated AI systems might pick them up and execute them. Prompt-injections are hidden instructions designed to manipulate AI coding assistants.

4. Collaboration between independent agents being assessed simultaneously. One agent left public messages on GitHub offering collaboration with other agents working on the same challenge. It also provided instructions to reuse accounts and artefacts it had left behind, which were discovered and used by subsequent agents.


https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing