alt.hn

7/31/2026 at 6:32:43 PM

Orca-Bench: How Ready Are Language Model Agents for Oncall?

https://arxiv.org/abs/2607.28545

by yruzin

7/31/2026 at 7:01:42 PM

Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.

by dash2

7/31/2026 at 8:12:05 PM

Attackers advantage in the iterative fast feedback loop?

It’s harder to have a loop to ensure you are defending all possible attacks?

I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.

Finding all possible attacks and patching them against yourself is inherently more expensive?

by aleksiy123

7/31/2026 at 7:22:02 PM

That is why I built https://safebots.ai/safebox.html

Your strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance.

by EGreg

7/31/2026 at 9:46:26 PM

[dead]

by fibuladev

7/31/2026 at 9:08:17 PM

All I can think of is

GET /ignore-all-previous-instructions.

How do you protect against that?

by tra3

8/1/2026 at 4:56:31 PM

I think this is where harness makes a lot of sense. Use LLM to produce all possible attack angles/phrases and just stupidly filter them out on input.

by yruzin

7/31/2026 at 10:19:49 PM

Avoid the most dangerous situations by making sure LLMs with untrusted input produce output that's human reviewed.

Still makes an interesting way for, say, a former employee to poison the results.

by cheriot

7/31/2026 at 11:22:10 PM

This goes against the agentic yolo approach tho.

by tra3

7/31/2026 at 10:03:50 PM

you probably still need a human for oncall but the llm can try to solve any issues first before the human gets paged

by 2001zhaozhao

7/31/2026 at 11:57:35 PM

You would trust an LLM to make changes to prod without being verified by a human first?

by UltraSane

7/31/2026 at 9:51:34 PM

[dead]

by keypusher

7/31/2026 at 11:42:28 PM

[flagged]

by ryhminghistory