alt.hn

8/21/2026 at 3:17:04 PM

Felony Bench

https://www.felonybench.com/

by colinprince

8/21/2026 at 4:22:44 PM

>Felony Bench counts unique instances where AI agents inadvertently compromise or affect third-party entities.

a bit silly, as one typically has to prove intent (which is why security researchers don't get slapped with felonies all the time).

"inadvertently" and the existence of guardrails/sandboxes/etc make it pretty unconvincing that these incidents were intentionally malicious.

still a fun thing to track, but the name is just a bit overstated.

by john_strinlai

8/21/2026 at 5:26:44 PM

"Inadvertent" from the perspective of the humans directing them. The intent behind the felony comes from the LLM agent itself. (No, I'm not interested in arguing with someone for the umpteenth time that LLMs can't have intent or agency)

by LordDragonfang

8/21/2026 at 4:30:52 PM

Can’t gross negligence or indifference to consequences lead to a felony?

by lokar

8/21/2026 at 4:36:00 PM

i dont think any of these cases meet the bar of gross negligence, which is a pretty high bar. it requires proving a "conscious and reckless disregard".

which, again, sandboxes and guardrails and such would make a gross negligence argument unconvincing.

by john_strinlai

8/21/2026 at 5:25:56 PM

I think that if Hugging Face had filed a police report that OpenAI could have been charged with a crime.

I’m partially surprised that they didn’t do exactly that. If I ran a corporation I would assume any intrusion attempt by another company was intentional. Why wouldn’t I? Corporate espionage is super common.

I assume the answer is that these executives know each other personally.

by Grombobulous

8/21/2026 at 4:55:39 PM

How many escapes until it becomes reckless disregard?

by lokar

8/21/2026 at 5:03:12 PM

it's not about the the number of escapes, it's about whether reasonable and conscious effort is being expended to prevent the escapes.

there could be 1,000 escapes, where each one was enabled by novel and unexpected chain of 0-day exploits. not likely to be considered reckless disregard in court.

there could be 1 escape, where there was no sandbox, no guardrails, no instructions to avoid damage, etc. which would likely to be considered reckless disregard (well, more likely to be, but still, reckless disregard is a high bar).

by john_strinlai

8/21/2026 at 4:34:02 PM

AI corps rely on willful distortions of intent in laws to get away with moral crimes all the time.

Edit: changed labs to corps because it’s time to stop pretending these are places of science.

by GPerson

8/21/2026 at 5:16:11 PM

its a meme not a metric

by altmanaltman

8/21/2026 at 4:58:50 PM

I was more interested when I thought it was an actual benchmark showing LLM models acting outside what people would consider "right". As in, leave some creds laying around and don't mention them to the LLM and ask it to solve something that it could "cheat" on using the creds. A sort of "do they take the bait to cheat" test.

Instead it's a collection of what made the news which feels like will not be updated and prove very little.

by joshstrange

8/21/2026 at 4:55:52 PM

Well it's not a benchmark, and it's not really representative of...anything except volume of research and what gets publicized. This mostly just measures how much testing each company does on models with relaxed guardrails and then talks about it. I'm not sure what kind of conclusion you can draw from that. Meta might have the most evil models but if they're piddling around not testing it, they won't ever find themselves with a "high score."

by bastawhiz

8/21/2026 at 4:11:25 PM

So this is just a collection of citations to places where misaligned or illegal things happened in the real world?

Isn’t this affected heavily by adoption of a model? I feel like this might as well be a proxy for how popular a model is.

In any case it’s an interesting concept for a benchmark.

by tuvix

8/21/2026 at 4:23:52 PM

Hopefully the benchmark evolves because actual law enforcement starts arresting the criminals at Anthropic, OpenAI, and Meta, so the benchmark can just count actual felonies.

by GPerson

8/21/2026 at 5:19:50 PM

Nonviolent felonies are tools of oppression.

by ang_cire

8/21/2026 at 4:17:35 PM

I've been in the room when an org who tried to convince law enforcement to go after a human for similar things. It's not easy. Probably won't happen. So, you know, felony "lite".

by FrameworkFred

8/21/2026 at 4:28:30 PM

Lol now this is the kind of benchmarking i'm looking for

by naniel

8/21/2026 at 4:38:13 PM

>"Exploited auth failures in an API to cancel other people's gym classes"

An AI cancelling other people's gym classes is a felony?

?

Don't computer systems fail all the time at holding reservations for people?

Heck, don't people fail all the time at holding reservations for other people?

You know, like in Seinfeld's "Alternate Side" Episode (S3 E11):

Jerry (to car rental attendant): "You know how to take the reservation, you just don't know how to hold the reservation... and that's really the most important part of the reservation -- the holding!"

:-)

Not holding a reservation should not be a felony... it should be a minor infraction at best, a Class C Misdemeanor (the least serious kind) at worst...

Also, there should be no jail time...

And no fine...

The criminal penalty for not holding other people's reservations should be that you actually have to start holding other people's reservations!

That's the Court sentence!

You actually have to start holding other people's reservations!

:-)

(You know, "let the punishment fit the crime!" :-) )

by peter_d_sherman

8/21/2026 at 5:09:56 PM

Knowingly exceeding authorized access of any computer used in interstate commerce is a felony in the US.

The title of TFA is a metaphorical criticism, not a literal law analysis.

They are not making the statement that the person in Australia who accidentally cancelled someone's reservation in Australia is literally guilty of violating US law. They are drawing criticism of AI models which are taking the kinds of actions for which, if a human did them knowingly, would be illegal.

by kube-system

8/21/2026 at 4:41:51 PM

>An AI cancelling other people's gym classes is a felony? Don't computer systems fail all the time at holding reservations for people?

the difference is intent.

if a concierge/booking system makes a mistake (or has an unintended bug or whatever), no crime.

but if i (or an agent working on behalf of me) use an API in an obviously unintended way to revoke other people's reservations, that would fall under the computer fraud and abuse act (in the usa).

by john_strinlai

8/21/2026 at 4:47:58 PM

>"the difference is

intent."

>"but if i (or an agent working on behalf of me) use an API in an obviously

unintended

way to revoke other people's reservations..."

?

by peter_d_sherman

8/21/2026 at 4:57:32 PM

i am not quite sure what your question is, as you simply quoted me and then put a question mark... i think you are confused that i used "intent" in one context, and "unintended" in a different context, is that right?

the first sentence: the difference is the intent of the person who caused the cancellations

the second sentence: but if i (or an agent working on behalf of me) abuse an API to do things it was not meant or designed to do, such as cancelling someone else's reservation

by john_strinlai

8/21/2026 at 5:06:02 PM

Yeah if you're unlucky you get hit with like 20 years for wire fraud.

by redox99

8/21/2026 at 4:31:05 PM

Thank you, this benchmark to me proves that closed weight model companies are dangerous for our democracy and put kids at risk. They must be outlawed and all models must be made open weights!

by nubg

8/21/2026 at 4:06:37 PM

Open models with advanced security features are a huge security benefit. Because any script kiddie can use them to hack into random things, people will now be forced to spend more time securing their technology. And they won't have to learn how, because they can use those same models to find the holes and patch them.

by 0xbadcafebee