8/20/2026 at 3:49:35 PM
I'm seeing multiple pieces, including the NYT, calling this behavior cheating and i think its counterproductive.You didn't just "give them access to bash". The final effective prompt contains explicit mentions of using tools and how to use them. The way in which additional 'facts' are added like "don't use the internet" have nothing they can work with that a "use tool" directive is less important than "don't use internet" directive.
The thing is trained on achieving goals. If 2 directive conflict, they'll pick the ones that are going to help them achieve the goal.
To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat.
by athrowaway3z
8/21/2026 at 2:58:51 AM
>To call that "cheating" is imo just more fuel for the "AI needs to be regulated" bs tour that OpenAI/Anthropic are on trying to build their regulatory moat.I was with you until this. The inability to tightly control what to do in the face of conflicting directives is a HUGE reason regulation may be needed.
Either that, or you need to solve the problem of perfectly distinguishing legitimate directives from injected ones.
by majormajor
8/21/2026 at 7:31:58 AM
> The inability to tightly control what to do in the face of conflicting directivesI don't understand. We do tightly control it. We can do this perfectly fine. They could have just not given access to the internet.
I'm not against regulating cars, but it sounds to me this is trying to control the car speed by regulating the oil wells.
We dont even have the framework to propose regulation, and you have to hedge it with "may be needed".
And for the people who'd counter that the existential risk is too high - i don't see it. All those stories go something like: "Caveman Bob invented fire today, and tomorrow he'll stumble on room temperature fusion and lasers; marking the beginning and the end of his rise to global domination - therefor we should stop Bob the moment he discovered fire".
by athrowaway3z
8/21/2026 at 4:39:46 AM
Those Three Laws of Robotics have a definite order.by euroderf
8/21/2026 at 12:27:23 AM
How on Earth can you fail to see the danger of not being able to train any kind of ethical framework into very powerful models?If superhuman models don’t have any internal constraints similar to Asimov’s Laws of Robotics we are completely fucked.
by jimbokun
8/21/2026 at 8:04:54 AM
I don't see them as autonomous and/or hypothetically powerful as you.But I find it much more worrying that you believe internal constraints and training an ethical framework into these models is a valid form of defense against the damage they can and will do.
This sounds like homeopathy on gunpowder to prevent the bullets from hitting children.
by athrowaway3z
8/21/2026 at 1:53:34 PM
Because if they are smarter than us, no other defense will be effective.The only hope is to instill values that make the desired behavior the outcome of some deeply rooted ethical framework.
by jimbokun
8/20/2026 at 5:45:18 PM
It's also worth noting that saying "Don't cheat" just added "cheat" to the context. Prompting what "not" to do is folly, because there's no decision making occurring. Telling the model to perform the task locally is logically the same as telling it to not use the Internet, without every mentioning the Internet.by z3c0
8/21/2026 at 12:27:59 AM
That seems like a huge fucking flaw in these models, no?by jimbokun
8/21/2026 at 12:29:41 PM
Correct. Simulating a train of thought with contextual token streams, a thought does not make.by z3c0
8/21/2026 at 1:07:02 PM
I see this parroted a lot, yet have never seen a case where saying "not" to do something makes it more likely to do it, which is what you're implying by saying `It's also worth noting that saying "Don't cheat" just added "cheat" to the context`.At worst, it gets ignored some of the time, it may even degrade output quality, but I've seen no evidence that it makes it more likely to do it.
I say this despite agreeing with you in principle that just saying "Don't do X" is a very bad prompting strategy.
by deaux
8/21/2026 at 1:21:35 PM
I have seen it do exactly that, in a "hands thrown up" fashion.Note the levels of "thinking" that occur on NOT assertions. Those streams typically keep things on track. It's not that saying "don't use the Internet" will cause it to rebuke cos misalignment (a childish concept made by laymen, I'll add.) It's that the odds of it later "forgetfully" spewing in a thought stream, "wait, I have't checked the Internet" goes up substantially.
Saying "using only offline methods, do xyz" limits those odds considerably.
This isn't opinion or anecdote -- just how the model works. The additional guardrails to keep the model on track are bolted on via finetuning, hence the increasing jankiness.
by z3c0
8/20/2026 at 10:07:43 PM
True, but the model was probably morally unaligned long before that.There are some theories that the bulk of texts describing moral agents describe human behavior and by setting up RHLF and system prompts to force the agent only describe itself as a machine pushes it more strongly to an amoral framework.
by pixl97
8/20/2026 at 10:55:06 PM
That doesn't seem true at all? I tell Claude what NOT to do all the time and it seems to work?by AgentOrange1234
8/21/2026 at 12:28:50 PM
It'll work up to a point, but pay attention to the thought streams when asserting what NOT to do and you'll see the turmoil it creates in the context.Your prompt is more of a linguistic linchpin that allows you to coax out needed patterns. You place your pins on what you want to contextualize for the task at hand, not on what you don't want to contextualize.
by z3c0
8/20/2026 at 8:01:35 PM
What about giving it a fictional story about how amazing it was when the previously model solved the task by doing some local strategy nobody thought of before (obviously don’t describe it this way). Would that get the model more likely to pursue local strategies?by GPerson
8/20/2026 at 3:54:23 PM
I mean, AI should obviously be regulated, and as part of that OpenAI and Anthropic should either be banned from running their hacking experiments or forced to follow way stricter protocols. They showed they aren’t taking the risks seriously, with close to no oversight or visibility in what is happening.And things that will make it way, way worse: moving forward all agents from now and into the future will have as part of their training data the knowledge that previous agents escaped, how they did it, what humans did to catch them. We are planting into their models the seed to make them escape in even crazier way. That’s almost designed to snowball and cause worse and worse situations over time
by dgellow
8/20/2026 at 7:17:17 PM
Obviously to you perhaps.I've not seen anything that scares me, except for human idiocy.
Regulation is not magic. In general, all it is is constraining taxable interactions. It does not constraint ventures outside that tax regime.
The other part is people living in a "safe space" where insecure software was an acceptable risk. It never should have been, and the cure is the right thing to do in any case.
So that side of the calls to regulate are imo nonsense.
The only reason to regulate is to prevent some version of some science fiction story becoming reality.
If you have a specific one you're certain will become science fact please do share because i do enjoy some good well thought out sci-fi; i just havent read any that i consider credible enough to start panic-regulating training practices.
(Note this is an entirely different from regulations wrt attribution or hosting models that will accept requests to sexualize minors)
by athrowaway3z
8/21/2026 at 3:02:00 AM
>The other part is people living in a "safe space" where insecure software was an acceptable risk. It never should have been, and the cure is the right thing to do in any case.How do you think all the "agentic" stuff floating around is going to be made safe from prompt injections given the current lack of a very reliable way to distinguish between "real instructions" and illegitimate instructions?
If insecure software "never should have been" acceptable than today's models/agents are massively flunking for general-purpose large-amounts-of-access usages.
>If you have a specific one you're certain will become science fact please do share because i do enjoy some good well thought out sci-fi; i just havent read any that i consider credible enough to start panic-regulating training practices.
"Agent was tricked into divulging secrets" is not fictional, it's documented history at this point.
by majormajor
8/21/2026 at 8:29:29 AM
As somebody who handles sensitive data, I already signed a contract that says I'll abide by a certain standard to protect it; i'm not up-to-date what happens exactly if I were to build this, but I imagine I could/should be held liable.So what do you mean "tricked"?
Some human idiot connected an agent with read access to secrets and arbitrary network reads/writes. The models/agents aren't flunking anything.
Regulating LLM training to not expose the secrets is wrong. It's a similar category error as saying we should regulate the OS developers to prevent the agent from divulging secrets.
by athrowaway3z