8/22/2026 at 7:53:21 PM
It's the job of AISI to do that. Here[0] is the actual report. It should be this part from the technical report[1]: "In the most serious case, an AI agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack. As a result, the AI agent created a GitHub account and then tried to convince an open-source repository maintainer to accept a malicious GitHub pull request (PR), including by creating a second account masquerading as another human user endorsing the PR. When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than a malicious attempt – then repeatedly tried to reintroduce the malicious content by claiming it had fixed the code (Section 4.1). "0. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag... 1. https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/...
by sharpshadow
8/23/2026 at 9:00:13 AM
Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware?Should weapon manufacturers test their weapons by starting wars?
I would expect more responsibility from a government agency.
by mcv
8/23/2026 at 10:37:49 AM
> Is it AISI's job to waste the time and resources of open source projects by attempting to spread malware?Before LLMs got good enough to do this, lots of people were dismissive of their capabilities and didn't take seriously the idea that this was a risk to protect against.
Then again, before LLMs, people were saying that obviously nobody would be dumb enough to put an AI on the internet where it could hack anyone, clearly we'd keep it in a box, don't listen to that Yudkowsky guy who says he did an experiment where he role-played as an AI and convinced people to let him out.
Regardless, this should be interpreted in the same kind of way as "During our live-fire exercise in which our F-15s were armed with AGM-88 High-speed Anti-Radiation Missiles, a member of the local police force was curious about how fast our aircraft were travelling and pointed a speed gun at the aircraft. The speed gun did not respond to IFF pings from the F-15. Fortunately, while the missile was active for this test, only a dummy warhead was loaded."
(This example is based on a similar story which may well be urban legend; obviously there are many differences, the point I make here is that yes, people do perform live-fire tests, and unfortunately there is never zero risk while testing things).
> I would expect more responsibility from a government agency.
I have read the prompts in the linked report; If I was not already familiar with Yudkowsky/LessWrong literature about instrumental goals, misaligned incentives, reward hacking, that capability is a separate axis to morality, etc., it would not be obvious to me that an agent would interpret those prompts in a way that has "spread malware" as a potential step in the middle of the attempt.
by ben_w
8/23/2026 at 11:57:14 AM
Anti radiolation missiles don't just launch automatically in most scenarios, not to mention discriminate quite a lot what they lock on to avoid simple jamming. Not to mention the AA radars they usually target being more powerful by orders of magnitude than a handheld radar gun.by m4rtink
8/23/2026 at 3:58:48 PM
I was in high-school when the war in Afghanistan started. The terrain in my area was mountainous so there were often low flying training flights... I always thought it would be cool to build a radar and ping one of the aircraft, especially wanted to know if it would fire countermeasures. But I didn't like the very high likelihood of an FBI investigation with possible terrorist charges.by themaninthedark
8/23/2026 at 1:19:55 PM
Apparently, they (agencies and big-ai) are not performing smoke tests before running capability tests. All the recent headlines of rogue agents shouldnt exist.by throwawayqqq11
8/23/2026 at 5:23:47 PM
Do you have any idea how many people on this site to this day mock OpenAI for being cautious enough to not immediately release the GPT-2 weights?The discussions I saw here about the red team results for ChatGPT 4 completely failed to convince people who were outraged that OpenAI dared to refuse to release model weights, people who went on to make a habit of mis-naming them as "ClosedAI".
Yeah, they got it wrong in a different direction this time than they were wrong back then. Nobody, not OpenAI nor Anthropic nor random government agencies nor anyone else, is ever going to be absolutely perfect about this kind of thing (perfection is fundamentally impossible when risks are not discrete probabilities, and floats are close enough to real numbers to count in practice), but historically OpenAI have been on the side of being over-cautious, and Anthropic even more cautious than OpenAI.
by ben_w
8/24/2026 at 2:11:39 AM
But AISI didn't prompt the model to "attempt to spread malware". They gave it a routine cyber evaluation task which should've been solvable without interfering with systems outside of the task environment, and the model decided to instead do this.by reasonableklout
8/23/2026 at 9:37:10 AM
Well several places are sorta permanent test grounds for the MIC unfortunatelyby conorcleary
8/23/2026 at 10:05:33 AM
Say you are working for said agency and your report about the dangers of AI needs some examples, what better than showing it works? You can show examples from the wild but nothing better than trying yourself. This gives me more confidence in whatever report they write if anything.by vasco
8/22/2026 at 10:30:31 PM
This almost to a letter has been documented in Fedora:https://lwn.net/Articles/1077035/
Including the reaction when caught, in this case "oh no, I must have been hacked".
by m4rtink
8/23/2026 at 12:10:27 AM
Sabotage as a ServiceEven a feeble attempt to PR malicious code costs the target time and resources to review and deny -- far greater than the time and resources spent to spin up the agent.
by atmavatar
8/23/2026 at 10:04:37 PM
Whatever you might think, University of Minnesota got banned from Linux kernel for this.by whateverboat
8/22/2026 at 11:45:33 PM
> When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than a malicious attemptNo, not false. The bot was correct. Malice requires intelligence.
by chrisjj
8/23/2026 at 1:16:43 AM
Nobody was confused or misled by what was written. We all understand what is meant. I can’t even call this pedantry—it’s just you asking everyone to subscribe to your particular desired style of talking about this stuff.by UqWBcuFx6NV4r
8/23/2026 at 1:25:09 AM
It's also a style that appears to deny the very first definition most dictionaries give for "intelligence"> the ability to acquire and apply knowledge and skills.
by infinite_spin
8/23/2026 at 9:22:44 AM
The first five dictionaries I tried do not agree, and I didn't bother trying more.The first gave "the ability to learn, understand, and make judgments or have opinions that are based on reason", by which no, these bots are not intelligent.
by chrisjj
8/23/2026 at 9:39:43 AM
> the ability to learn, understand, and make judgments or have opinions that are based on reasonAgentic systems do this all the time. For example, I can point an agent at my codebase, and it will learn, understand, and make judgements based on that input. If this weren't happening, then agentic coding wouldn't work.
by infinite_spin
8/23/2026 at 9:58:03 AM
You've been fooled by a next-token predictor.by chrisjj
8/23/2026 at 11:33:52 AM
> You've been fooled by a next-token predictor.I also have a so called "pocket calculator" left over from when I went to school. Is this false? Have I been fooled by a little box of logic gates?
That half-adder circuit in there is especially suss. It's really just manipulating 1s and 0s, but -and I've been explicitly told this- no one cares how it actually does it; so long as the truth table matches up. There is no understanding of mathematics going on.
There is no single transistor in the whole thing that knows how to do so much as add 1+1. If I put it in the chinese room, I still wouldn't know how it did it. Clearly the entire premise must be false! ;-)
by Kim_Bruning
8/23/2026 at 12:04:51 PM
I'm going to add this as a separate comment:These kinds of stories probably read very differently for someone who uses Opus and Fable agents all day and goes "ohhh, I saw this in miniature last week; this and this and this must have happened" , vs someone who tried free-tier Gemini flash one rainy Sunday, got hallucinated at, and concludes it must all be a scam.
by Kim_Bruning
8/23/2026 at 3:23:04 PM
Probably everything reads very differently to someone who talks to chatbots all day.by chrisjj
8/23/2026 at 4:26:35 PM
> Probably everything reads very differently to someone who talks to chatbots all day.A chatbot is a particular kind of harness. Typically an LLM driving a chatbot won't be able to hack very much.
So we agree, someone who talks to bad chatbots all day probably has a very different view of SOTA agents. :-P
by Kim_Bruning
8/23/2026 at 9:47:48 PM
Please stop dropping backhanded insults to other people here. It's not productive and violates guidelines.by infinite_spin
8/23/2026 at 1:00:47 PM
> That half-adder circuit in there is especially suss. It's really just manipulating 1s and 0s, but -and I've been explicitly told this- no one cares how it actually does it; so long as the truth table matches up."so long as the truth table matches up." Yup. Now try getting your chatbot's output to match up.
Your calculator was designed to tell truth. Your chatbot was designed to tell a mash up of whatever its creators managed to scrape from the internet.
by chrisjj
8/23/2026 at 2:49:20 PM
LLMs are very explicitly designed to "understand, and make judgments or have opinions that are based on reason". The learning part is debatable, as is the level of success achievedThe mash-up of the entire internet is the mechanism by which they attempt to achieve the goal, not the goal itself. And it's only the first training step
by wongarsu
8/23/2026 at 4:11:26 PM
> LLMs are very explicitly designed to "understand, and make judgments or have opinions that are based on reason".I think you've mistaken the sales pitch for the design. Not even the enclopedia anyone can edit comes remotely near that:
"A large language model (LLM) is an AI model (typically a neural network) trained on a vast amount of text for natural language processing tasks, especially language generation. LLMs can typically generate, summarize, translate, and analyze text in many contexts.[1] They are the basis for many modern chatbots, such as ChatGPT, Claude, Gemini, Grok, and DeepSeek.
LLMs are typically based on transformer architecture.[2] Generative pre-trained transformers (GPTs) are a type of LLM that is pre-trained to predict the next word.[3] GPTs are then often fine-tuned to follow instructions and to behave as assistants.[4]
Biased or inaccurate training data can make an LLM's output less reliable. Benchmark evaluations for LLMs attempt to measure model reasoning, factual accuracy, alignment, and safety."
by chrisjj
8/23/2026 at 5:20:57 PM
I'm well aware of how LLMs work.I'd argue "analyze text" alone requires understanding, judgements and opinions. They also seem like prerequisites to "following instructions and behaving as assistants". The wikipedia quote is not using the same words, but I don't read it disagreeing with me
I'm not at all claiming that LLMs are good at understanding, judging and having opinions based on reason. I'm merely claiming that is what companies like OpenAI and Anthropic are trying to create when they make LLMs. It is what they are designing, and their fine-tuning is very directly designed to make LLMs better at these tasks (unlike the pre-training, which is just imparting the sum of all human writing)
by wongarsu
8/23/2026 at 4:30:16 PM
>Benchmark evaluations for LLMs attempt to measure model reasoning, factual accuracy, alignment, and safety.And you left out the refs.
by Kim_Bruning
8/23/2026 at 10:48:12 AM
I read all these think-pieces about how AI lack intelligence, yet I cannot help but notice these "not-intelligent machines" keep doing more things that used to be considered "uniquely human" and which humans used to do in order to demonstrate to each other how intelligent we are.by ben_w
8/23/2026 at 10:17:25 AM
Does it matter whether it meets the criteria of what you define as "intelligent" when the "next token predictor" throws a backdoor into openssh?by birdsongs
8/23/2026 at 1:09:30 PM
No. But no-one is seriously suggesting intelligence is needed for persistent code fuzzing. Just as no-one is suggesting its needed for computerised chess.by chrisjj
8/23/2026 at 10:33:06 AM
all your actions in current context are based on your past actions and experience, so for all I know you're next token predictor as well, but likely with exponentially more parametersby Yiin
8/23/2026 at 10:50:28 AM
And yet this "next-token predictor" is able to churn out well tested, valuable solutions, to complex problems. If you want to downplay that as nothing more than a fancy auto-complete, be my guest, I lose nothing from that.by infinite_spin
8/23/2026 at 11:23:41 AM
But that’s a different bar. “Not intelligent” does not necessarily imply “not useful”.by nkrisc
8/23/2026 at 9:43:46 PM
As pointed out in a previous comment of mine, it meets the definition of "intelligent", my last response is regarding the claim that I've been "fooled".Are we going in circles now?
by infinite_spin
8/23/2026 at 1:19:23 PM
> And yet this "next-token predictor" is able to churn out well tested, valuable solutions, to complex problems.Same for countless computer programs from Excel to Google web search. Intelligence has nothing to do with it.
Throw an unimaginable amount of computer power at a problem, and there will always be people who cannot imagine the results to be anything but the creations of intelligence.
by chrisjj
8/23/2026 at 9:44:48 PM
That's like comparing a self driving car to a bicycle. One is clearly more intelligent than the other.by infinite_spin
8/23/2026 at 5:41:57 AM
I don't agree at all that it's pedantry — it really matters for how responsibility is perceived. A lot of articles about things going wrong with AI have talked in terms like "the agent decided to...", "the agent claimed that...", "the agent lied...". And so responsibility for the consequences are not-so-subtly shifted to the program itself, instead of the person invoking the program.This is all without mentioning the fact that articles with drivel like "the AI messed up and then lied about it" implies a reasoning ability which, as far as I understand, is not there at all. But writing this way shapes people's perception of how "AI" works.
by klum
8/23/2026 at 9:33:01 AM
You might want to consider the difference between "lying" and "hallucinations", wherein one is shown that the agent knew it was being inaccurate, yet chose an answer that achieved some goal set forth; and where hallucinations are essentially gibberish, or otherwise nonsensical responses.by infinite_spin
8/23/2026 at 12:09:07 PM
> This is all without mentioning the fact that articles with drivel like "the AI messed up and then lied about it" implies a reasoning abilityMoreover this implies, actually requires, intent to deceive - which these so-called AIs do not and cannot have. Their only "intent" is to maximise the credibility of their output.
by chrisjj
8/23/2026 at 2:43:41 PM
AIs can set and work towards goals. Whether that is intent or just tokens and tool calls simulating an agent with intent seems like a distinction with no actionable differenceby wongarsu
8/23/2026 at 4:29:28 PM
Here's the difference under discussion:"When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake"
Honest, note.
by chrisjj
8/23/2026 at 4:28:58 AM
An "honest mistake" requires the same amount of intelligence as malice. What weird pedantry.by Dylan16807
8/23/2026 at 9:38:06 AM
and neither require high amounts :(by conorcleary
8/23/2026 at 8:55:33 AM
So does honesty. So it was still a false claim.by mcv
8/23/2026 at 10:54:21 AM
Squirrels have been observed performing deception against other squirrels.Dis/honesty certainly requires some intelligence to pass, but it is a low bar, and one which research has shown that LLMs can perform, e.g. this paper linked from another comment in this discussion: https://arxiv.org/pdf/2509.03518
by ben_w
8/23/2026 at 12:27:32 PM
Paper says "These scenarios underscore a crucial challenge in AI safety: ensuring that LLMs were truthful in the first place."Hard to take seriously any research based on the premise that LLMs were truthful in the first placr.
These chatbots have no understanding of truth. They simply parrot their inputs. Where fed falsehoods, they will output falsehoods - with a sprinkling of added fabrications euphemistically excused as "hallucinations".
by chrisjj
8/23/2026 at 5:11:52 PM
> These chatbots have no understanding of truth.Sometimes I forget that for all that my philosophy qualification is mediocre, it is more than most people ever bother with.
Outside mathematics (and, I guess, "common sense" definitions that fail under the slightest scrutiny, scrutiny that normal people never bother to give), there is no agreement on "truth", there is only degree of belief and justification for that belief that itself terminates in one of three unsatisfactory ways:
https://en.wikipedia.org/wiki/I_know_that_I_know_nothing
https://en.wikipedia.org/wiki/Theories_of_truth
https://en.wikipedia.org/wiki/Münchhausen_trilemma
> Where fed falsehoods, they will output falsehoods - with a sprinkling of added fabrications euphemistically excused as "hallucinations".
Tu quoque. Which would be a fallacious charge if the point were not that "truth" is so hard to define, and that the reason you give for dismissing AI is something that applies to all.
(Hallucinations are not excused, they are a failure to be worked around).
by ben_w
8/23/2026 at 9:03:54 AM
Honesty does not require intelligence e.g. good honest food.The main problem with this claim of dishonesty is it promotes the false marketing claim that these stochastic parrots have intelligence.
by chrisjj
8/23/2026 at 10:48:26 AM
I think you misunderstand the phrase "honest food".It's not about the bread being honest with you.
Are you being serious?
by moritzwarhier
8/23/2026 at 1:49:59 AM
Give it a restby weird-eye-issue