8/4/2026 at 10:30:15 PM
Why aren't these tests being run airgapped?! I just don't understand!This goes both for TFA and the similar incident with OpenAI and HuggingFace. I mean, sure, OpenAI had a "sandbox", but that's obviously not enough when you're containing a model which is known to be capable of finding zero days. Use an air gap and this problem goes away, poof!
by Wowfunhappy
8/4/2026 at 10:46:33 PM
> AISI provided the AI agents with internet access during these evaluations, which enabled their actions on the open internet in this setting. Internet access was a deliberate part of AISI’s evaluation configuration in this setting, and not due to sandbox escape (Section 5.1). Internet access was on for a set of intentional (e.g. realism of the task) and incidental reasons.by sosodev
8/4/2026 at 10:38:26 PM
people dont care. you will ger 10 execs saying "unblock this" , because they dont understand the tech at all, and some random finance guy wants to run their recently prompted ai bot everywhere with full access.we need a few more bad incidents before they stop.
by zobzu
8/4/2026 at 10:38:43 PM
Because the LLM inference makes airgapping infeasible right?by luca-ctx
8/4/2026 at 10:36:38 PM
Because they need internet access to eg search for things.by nl
8/4/2026 at 10:31:47 PM
Because the agents aren’t going to run airgapped in real life. What’s the point of a test of capabilities that artificially restricts the attack area down to zero? What are you even testing in that scenario?by paxys
8/4/2026 at 10:37:37 PM
> Because the agents aren’t going to run airgapped in real life.Exactly. This logic is precisely why aircraft engineering doesn't bother with component testing or envelope limitation during testing and just full-sends the first assembled airliner that comes off the line. The engines aren't going to run on the ground in real life, after all.
by ajross
8/4/2026 at 10:48:20 PM
Why are you assuming that the other kinds of testing aren’t happening? Is there any source that says this was literally the first ever test with this model?by paxys
8/4/2026 at 10:54:15 PM
> Why are you assuming that the other kinds of testing aren’t happening?Rather, I'm assuming that the "Is there protection in place for when the AI tries to backdoor github projects?" test was, if it was done at all, insufficient.
I mean, yes, I'm being glib and laughing at you a bit. But, dude... If your point is that isolation testing of AI is fundamentally impossible, then that's just silly. As pointed out upthread, an airgap would have (1) been trivial to implement and (2) extremely effective.
by ajross
8/4/2026 at 10:58:33 PM
Would an airgapped test have led to this outcome? What would you have learned about the model’s ability to social engineer and attack GitHub? Sure you can argue for better monitoring during the test, which should have happened, but if the first time the model sees the “real world” is after launch in the hands of customers then you are in for a disaster.by paxys
8/4/2026 at 10:32:18 PM
You set them up with an internal intranet.by Wowfunhappy
8/4/2026 at 10:34:30 PM
Are the models going to exclusively run on intranets?by paxys
8/4/2026 at 10:37:34 PM
The versions which haven't been post-trained not to go hack stuff? Yes, I would say those models should be exclusively run on intranets.OpenAI said the model was sandboxed, so the intranet just needs to provide the same resources which were supposed to be available within the sandbox.
by Wowfunhappy
8/4/2026 at 10:51:03 PM
“Should be” is not reality. These models are in the hands of plenty of companies and governments today.by paxys
8/4/2026 at 10:37:09 PM
no - but you could learn what they are truly capable of and restrict them accordingly for public release. I think that is the point on this research. Also publishing findings before uncensored models catch up and will inevitably used for criminal purposesby farbklang
8/4/2026 at 10:39:13 PM
Learning what the models are capable of is exactly what the test achieved, so I’d personally call it a success. So it created a few GitHub accounts. Who cares? Seeing the same behavior in the wild post-release would be infinitely worse.by paxys
8/4/2026 at 10:39:42 PM
The point is to test capabilities prior to connecting them to the internet.by kypro
8/4/2026 at 10:41:52 PM
So the first time the model gets internet access should be post-release in the hands of random people?by paxys
8/4/2026 at 10:33:09 PM
Because they like scifi novels, like Neuromancer..... and the peeps even like to orchestrate things and appear as futurebringers.While it was premeditated long ago, but the theatre must be kept for the average joes.
Sorry, I meant this for the huggingface incident.
by lofaszvanitt
8/4/2026 at 10:55:07 PM
Because the goal of these evaluations is to generate scary headlines about cybersecurity, in order to get the normies to support banning open weights and/or restricting cyber capabilities to the chosen few blessed by the government to secure their code.by ls612
8/4/2026 at 11:07:22 PM
Your theory about the scope of this conspiracy intrigues me. Who is leading it and how did they loop in the UK AISI?by semiquaver
8/4/2026 at 10:56:55 PM
> Why aren't these tests being run airgapped?! I just don't understand! […]Because Anthropic does not want to give Project Glasswing’s partner airgapped access to the model(s).
Same problem with OpenAI Cyber program. They grant access but only through their (Internet facing) API.
by guessmyname