alt.hn

7/28/2026 at 8:52:55 PM

Codex Security

https://github.com/openai/codex-security

by bakigul

7/28/2026 at 9:30:15 PM

Hey HN, Michael here, co-founder of Promptfoo and one of the people working on the Codex Security CLI at OpenAI.

Thanks for checking this out and for flagging the auth issues. We just open-sourced it, and there's still plenty for us to improve. Expect the product to evolve quickly.

If you try it, I'd really appreciate hearing what works well and what you think we should improve. Happy to answer questions here.

CLI docs: https://learn.chatgpt.com/docs/security/cli

EDIT: If you'd like to help make this better, we're hiring: https://openai.com/careers/full-stack-software-engineer-cybe...

by dangelosaurus

7/28/2026 at 11:11:14 PM

> Have experience shipping production full-stack products across modern web frontends and backend services.

I'm amazed that the requirements are so low (or at least this vague) for jobs at companies like these.

Has anyone else had the experience of going to an interview and feeling like you were never asked any qualifying questions?

All the questions were easy, your answers were straightforward, you "got them right", but then were not chosen?

I find on the other side, they're also left with dozens of people who "passed" and then it comes down to a pretty arbitrary decision on who gets hired (if we are talking external, no referral, etc.)

I wonder if they can make job descriptions highly specific to filter the shortlist faster and more effectively (to actually get a shortlist).

Anyway end rant. Cool job, hope you fill it.

by orangelimesoda

7/29/2026 at 4:31:46 AM

I interviewed with them and had a great experience despite it ultimately not being a good fit due to communication issues and a bad flu coming on in the final interview - they, and Michael especially, was very kind and understanding.

by J_Shelby_J

7/28/2026 at 11:30:55 PM

This sounds like it’s just an app development role, not a security analysis position, so I’m not sure what your complaint is.

by jameshart

7/29/2026 at 3:10:41 AM

> just an app development role

Maybe this is part of the problem

by orangelimesoda

7/29/2026 at 1:21:02 AM

Professional app development requires an understanding of security.

by solid_fuel

7/29/2026 at 1:35:18 AM

I think the evidence contradicts you at this point

by corndoge

7/29/2026 at 5:54:31 AM

That's quite an indictment of the common app development practices. (Not that I disagree with your point...)

And security is hard. Because it is by definition off the happy path, it is quite often at odds with MVPs and rapid release cycles. Then you add all the ways the users can use your product to attack/abuse others.

Any non-hobbyist app development does indeed require at least a decent understanding of security.

by bostik

7/29/2026 at 3:17:39 AM

So you don't need to know anything about credit cards for someone to enter it into an app? I don't get the take (is it bad sarcasm?)

by orangelimesoda

7/29/2026 at 3:27:59 AM

To be fair, for the vast majority of cases, no you don't.

It is extremely rare for companies to roll their own payment processing anymore, or even handle PCI scope at all.

by devmor

7/29/2026 at 3:31:50 AM

Credit card entry, I think you should know a few basics like don't put it in MongoDB?? Or nah? It's just like any other user data?

How about a background check - can anyone take a user-entered DL and randomly Google stuff to see what they find?

Can I store your SSN in plain text in a text file? Why not?

The user had to upload their ID for IDV but I use Vercel. I guess I have to put it on S3. What should the bucket policy be for all these driver's license photos - there are so many???

by orangelimesoda

7/29/2026 at 7:39:05 AM

> Credit card entry, I think you should know a few basics like don't put it in MongoDB?? Or nah? It's just like any other user data?

> How about a background check - can anyone take a user-entered DL and randomly Google stuff to see what they find?

> Can I store your SSN in plain text in a text file? Why not?

You wouldn't be touching any of those unless you work for a handful of providers where that's their whole business. Usually you add the dependency, use their widget, and that's it.

by lmm

7/29/2026 at 3:37:07 AM

> Credit card entry

If you're handling credit card numbers yourself, you're in a shrinking subset of developer roles. I've spent the last 7 years of my career working at payment processors, so I do handle that stuff, but the majority of my industry has built an infrastructure that makes it so most developers don't have to think about that.

To your other examples, not everyone on the team needs to know these things up front. Someone in the review process does, and eventually that knowledge gets disseminated and more people know it to carry it forward in their career.

by devmor

7/29/2026 at 3:45:56 AM

[dead]

by redlimetea

7/29/2026 at 3:52:33 AM

To your last question about bucket policies, clearly you just need to scope it down: arn:aws:sts::*:assumed-role/trustme*/*

(indeed that first wildcard means any account)

by oenton

7/29/2026 at 5:16:16 PM

Good thing they asked for people with experience of professional app development then.

by jameshart

7/29/2026 at 2:44:40 AM

Not having clear, objective criteria enables arbitrary decisions (not against OP, I mean in general.. and in general I dislike this pattern a lot). On the extreme other end of the spectrum would be 100% objective criteria, and companies being forced to pick a random applicant that matches them. If they want only the best, they have to have high expectations, but be able to actually define them. You say "cultural fit", I say "corruption", let's blow this whole joint.

by customguy

7/29/2026 at 2:47:44 AM

> You say "cultural fit", I say "corruption", let's blow this whole joint.

It's not necessarily corruption. If there's no conflict between principal and agent, it's fine.

Just like when I sent my butler to go and buy a bottle of wine, he can make arbitrary choices, but that doesn't mean he's corrupt or going against my wishes. I trust his judgement, and since it's a repeated game our incentives are aligned.

by eru

7/29/2026 at 2:45:17 PM

If the company got any investment, than the principal has no interest in "cultural fit" and it's always corruption.

by marcosdumay

7/29/2026 at 3:32:03 AM

It’s not corruption to simply hire the people you subjectively feel a preference for working for, instead of objective criteria. It’s not public tax funds that fuel salaries, it’s your own money. You get to spend it how you like.

Yes, many jurisdictions have outlawed arbitrary discrimination against protected classes (eg race), which is an entirely different matter, and not what we are discussing here.

by sneak

7/29/2026 at 7:33:47 AM

Yeah, and it's not necessarily mobbing to only tell people you really like about your party. But where it occurs, people never admit it to themselves and justify it in such a way, so that justification is meaningless. All wars of aggression are called a defensive emergency measure. Hundreds and thousands and rarely would anyone say "we'll take this because we can and you're helpless". No matter how glaringly obvious it is, it's never admitted.

In the same way, I can acccept "cultural fit" as a summary of things a person can describe, sure. But I think more often than not it's just a thought-terminating cliché. It can also just mean "I cannot verbalize my reasons and/or don't want to admit to them".

You can say if something fits only if you either can describe both sides in sufficient detail and where it wouldn't fit, e.g. a plug and a socket. But if it's dark, you barely see anything, and just have a "hunch", then "fit" doesn't even apply. It's like telling someone you won't let them through a door because they wouldn't "fit" anyway -- okay, so let them try, if they actually won't fit you don't need to read tea leaves and gate keep based on that.

Preferring to go with a more safe and familiar and obvious candidate, fine. But don't pretend it's because the others won't "fit".

Culture, in so far as it deserves the name, shapes the people exposed to or in it, as well as the other way around. E.g. if only people who fit the culture can work at a company, no company can exist in the first place, because for there to be a culture there need to be people there. So that leaves setting the culture in stone after it grew to a certain size, and only looking for more of the same, which also isn't great, but at least still honest.

If a culture is so brittle it cannot integrate people who aren't already a product of it, that may be a legitimate choice of the company, but my assessment to find that lame is also valid.

I feel the same way about immigration troubles, tangentially. We moan because people we don't actively try to get to know don't care for our rules which we don't enforce in a confident, but respectful manner. We basically require sterile, bland input because we have no immune system worth speaking of and no way to process and refine what comes in.

by customguy

7/29/2026 at 8:35:37 AM

I agree the phrase "culture fit" is weasel-wordy, but...

> You can say if something fits only if you either can describe both sides in sufficient detail and where it wouldn't fit

Many human dynamics, including sexual attraction, love, and even just who will be fun or easy to work with, are dynamics we don't fully understand, and cannot fully specify. In all these cases "I'll know when it see it" is perfectly reasonable, and need not be hiding an untoward motive. Which is not say, ofc, that it can't be hiding such a motive. That happens too.

by jonahx

7/29/2026 at 8:33:51 AM

All fair points and I guess I would dislike it more if it hadn’t benefitted me more. In general, if I can get face-to-face with a human then I have a huge advantage, and any barriers to getting to that point are a net negative to me.

At the start of my career I was under-credentialed, and had to rely on lax requirements to get myself in front of people. I might never have gotten anywhere if my first few employers hadn’t been willing to overlook a lack of degree or commercial (rather than open-source) experience, directly as a result of a funnel with a wide entrance.

by petesergeant

7/29/2026 at 6:13:11 AM

It depends whether your goal is to hire for specific knowledge (hence specific questions) or for overall mindset and abilities (hence broader questions where you are able to extract the way a person thinks)

by gbalduzzi

7/28/2026 at 11:47:41 PM

If you didn't already know, jobs at highly competitive companies tend to have vague job requirements because they expect to be able to apply your raw intelligence to changing demands quickly. There's no point being hyper-specific about the exact software packages because that's not what they want. What they want is someone who, after talking to an interviewer for 30 minutes, leaves them with the thought "Wow, this person can do anything we need of them. They can probably tell us what we need too and take ownership of large projects. Hire!"

by pertymcpert

7/29/2026 at 12:17:43 AM

Agree. Being smart, competent, and high in conscientiousness is more important than any highly specific “qualification”. It’s not about checking a bunch of boxes. If you have a track record of getting shit done, you’ll have something to contribute.

by moscoe

7/29/2026 at 3:15:06 AM

> There's no point being hyper-specific about the exact software packages because that's not what they want

Okay. This makes it sound like they're more sophisticated but it seems more like they are less sophisticated, less specific, and a lot more vague in the job descriptions they themselves create.

If you look at any technical role, game dev or something where people are building important things at scale - there are a lot of specifics. Libraries, methodologies, where if you didn't know them you are nowhere near a fit.

I'm just wondering. It's OpenAI. Surely there is some domain-specific something beyond "has experience shipping front-end and back-end services" since that includes basically everyone.

It makes this job look like a Starbucks role.

by orangelimesoda

7/29/2026 at 3:32:37 AM

I see where you're coming from, but as someone who regularly interviews engineers, I don't care about specific tech stacks when evaluating a candidate very much either. I can only think of two positions I've worked in where such a thing really mattered.

A good engineer can adapt and catch up without a lot of lead time. For a contractor, I'd be much more specific - but for someone who's going to join my team? I'm looking for a candidate that can demonstrate their problem solving ability, creative thinking and communication skills.

Other than having some kind of experience in the general domain we work in, those "soft skills" are far harder to find than specific tech experience.

by devmor

7/29/2026 at 2:10:10 PM

Yeah, no. If you are digging deep into database internals, query optimization and schema design, it helps to have someone with some experience in the domain. Otherwise your team will spend a few years learning from first principles.

by a34729t

7/29/2026 at 4:12:04 AM

This is insane to me.

Not every engineering job is entry level or as simple as most fullstack crud. Deeper into industry you find highly specific well defined positions for a given domain. Soft skills matter more the higher the ladder but id take a killer senior who can be difficult over a team of mediocre staff engineers.

by rustystump

7/29/2026 at 3:39:32 AM

[dead]

by berrylimetea

7/29/2026 at 1:52:01 AM

This. Being too specific on the stack requirements is a red flag imo.

by _superposition_

7/29/2026 at 2:19:44 PM

> raw intelligence

I'm not sure how this would apply. Are you implying that if the company operates on Python, you can hire someone with great "raw intelligence" who have only developed C++ all their life, and they can start contributing on day 1?

You need to clearly list what the position entails, otherwise you're wasting time.

by swat535

7/29/2026 at 3:15:46 PM

They are listing what the position entails. In this case it's a full-stack web dev role. "only c++ all their life" is more or less not possible if you have experience doing that

Likewise from another job post: "Have a strong background in kernel-level systems". Not possible if your whole resume is building web apps, may be possible for someone that has only used c++ professionally.

by yesb

7/29/2026 at 2:40:37 AM

I agree, I recently interviewed at Synthiolabs, and they asked me two questions: one about RAG and two about graph RAG. They rejected me even though I answered correctly, and the interviewer was also a college kid.

by hacket04

7/28/2026 at 10:29:50 PM

This looks great, thanks for open-sourcing it!

How does it deal with the current guardrails 5.6 Sol has on finding vulnerabilities? When I use it in the Codex app it would sometimes say it found a vulnerability, but it cannot tell me what it is.

by vladoh

7/29/2026 at 12:00:25 AM

Thanks! You've run into a real limitation: the CLI doesn't bypass the model's cybersecurity guardrails. If GPT-5.6 Sol finds a vulnerability but refuses to explain it, switching from the Codex app to the CLI won't automatically fix that.

For authorized defensive work, Trusted Access for Cyber (TAC1/Daybreak) can reduce refusals depending on the model and the account or organization where access is provisioned. It isn't a blanket bypass.

If you're an open-source maintainer, you can apply for conditional Codex Security access here:

https://openai.com/form/codex-for-oss/

For enterprise teams, the public Daybreak onboarding guide is here:

https://help.openai.com/en/articles/20001261-enterprise-dayb...

If you have an example of "found a vulnerability but won't tell me what it is," I'd love to take a look too. You can send it to use with /feedback (or message me).

by dangelosaurus

7/29/2026 at 9:24:45 AM

> If you're an open-source maintainer, you can apply for conditional Codex Security access here:

> https://openai.com/form/codex-for-oss/

Hey, Lead maintainer of vim here. Applied twice already never heard anything back. This is a frustrating experience!

by chrisbra80

7/29/2026 at 2:57:28 AM

>> For authorized defensive work, Trusted Access for Cyber (TAC1/Daybreak) can reduce refusals

Or perhaps a better option is to use something like Kimi K3 and cancel the GPT subscription altogether.

by theplumber

7/29/2026 at 5:44:11 AM

Or try Grok, 4.5 seems pretty capable, should be close to K3 in many coding tasks. I use it for code review of what other "stronger" models shit out (like Sol) and it constantly finds even pretty big bugs or just not robust enough solutions (Sol tends to overengineer, yes, but I'm not so sure it overengineers the right parts, so far my experience woth it has been mid. Except it understanding my drawings and collages and it being capable of far better frontend/design dev than 5.4 or even 5.5 was).

by Culonavirus

7/29/2026 at 10:23:17 AM

Sounds like that's the only solution. I'm so sick of this safety nonsense I was going to switch from Anthropic to OpenAI because of it. I'm so disappointed to see it's just more of the same.

Model finds a vulnerability in your code but "refuses" to tell you. Words can hardly express the sheer absurdity of it.

by matheusmoreira

7/29/2026 at 1:45:22 PM

It’s like they want to squeeze more money from you with the cyber crap…reminds me of all the DRM stuff around music distribution…

by theplumber

7/29/2026 at 8:52:51 AM

> GPT-5.6 Sol finds a vulnerability but refuses to explain it

I think it would be a good practice to refund the session cost in that case. Otherwise a customer just spent some money in order to get exactly nothing.

by egorfine

7/28/2026 at 10:48:54 PM

I tried it, it started a scan but stopped after hitting the rate-limit of my account. It gave up after just a minute of retrying (rate limits are tokens per minute, so... :P).

It said "Partial output was kept at <...>", but I dont see a obvious way of picking it up in a new scan? (The failed run cost me ~$13)

by Quai

7/28/2026 at 11:28:41 PM

Yeah, you're right. A per-minute rate limit shouldn't kill a scan after a minute, and "partial output was kept" makes it sound like you can pick up where you left off. You can't yet, unfortunately. --max-cost can limit estimated spend, but we still need proper retries and resume. Sorry you spent $13 finding that out. Please send me an email and I'll help make it right.

by dangelosaurus

7/29/2026 at 2:08:20 PM

Off topic: Just some thx and kudos to you guys. I used Promptpoo at the beginning of the year - it was exactly, what I needed, very much still a niche thing hardly anyone was using.

I totally missed the acquisition - but well deserved. I am currently re-evaluating PF again for my upcoming project, and happy to see that it is more than simply thriving.

by _the_inflator

7/29/2026 at 2:42:06 PM

unfortunate typo lol

by MattRix

7/28/2026 at 10:06:47 PM

Does it require hitting OpenAI's APIs or can one also stand up a local OpenAI compatible LLM endpoint?

by strictnein

7/28/2026 at 10:10:41 PM

We are actively working on officially supporting this. Because it's open source it is pretty easy to point a coding agent at it now and switch out the model.

by dangelosaurus

7/28/2026 at 10:20:29 PM

Exciting! Is there any open GitHub issue we can track?

by ignoramous

7/28/2026 at 9:49:45 PM

When would I use this over the plugin in codex? Which I think can be invoked from cli as well

by gizmodo59

7/28/2026 at 9:56:27 PM

The plugin, including when invoked through the Codex CLI, is great for scanning the repo you're currently working in. The standalone Security CLI/SDK uses the same scanner, but is built for running security across many repos over time: org-wide scans, historical results, deduplication, false-positive tracking, budget controls, and CI integration.

We've been talking to hundreds of engineering and security teams, and their feedback is shaping what we build.

Like Promptfoo, our goal is practical tooling that fits into the workflows teams already have.

by dangelosaurus

7/28/2026 at 9:37:17 PM

Been watching your progress for a while, glad OpenAI have looked after you and the team and you still get to ship!

by robotswantdata

7/29/2026 at 12:24:14 AM

Thank you, that means a lot. Being able to keep building practical, open-source security tooling was important to us.

Really glad we got to ship this, and there's still a lot we want to improve in Codex Security and in Promptfoo!

by dangelosaurus

7/28/2026 at 10:41:32 PM

How does it fare against its own codebase?

by 6thbit

7/29/2026 at 7:27:54 AM

Hi! Any chance you may have tangential positions opening in Zürich?

by noname120

7/29/2026 at 12:43:56 AM

Why does this need an entirely separate repo instead of being a feature in the existing Codex project?

by waterTanuki

7/28/2026 at 10:13:23 PM

> co-founder of Promptfoo and one of the people working on the Codex Security CLI at OpenAI.

> Thanks for checking this out and for flagging the auth issues.

Offtopic, but this right here is why I don't believe any marketing around "great amazing models that one-shot everything and programmers are no longer needed".

You just have to look at what these labs routinely produce, and their own products.

Edit to respond to @simonw whose comment I saw before he retracted it ;)

This comment is tied directly to consistent continuous claims by the LLM labs. Their own products disprove their own claims, and it would indeed be nice if fewer people believed them :)

by troupo

7/29/2026 at 3:54:56 AM

[dead]

by imrozim

7/29/2026 at 7:51:56 AM

Hi! Any remote internship for a high schooler? lol

by sudo_cowsay

7/29/2026 at 4:39:31 AM

Hey looks cool. I tried to run this on a small oss library and here's what happened:

  $ codex-security scan .
  [00:00] Preparing scan
  [00:00] Authentication: stored Codex credentials.
  [00:01] Preparing scan
  [00:42] Running scan
  [00:42] Preflight: worker delegation supported (up to 8 worker slots).
  [41:03] Running scan
  codex-security: This content was flagged for possible cybersecurity risk. If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program: https://chatgpt.com/cyber
  codex-security: Partial output was kept at /Users/ryan/.codex/state/plugins/codex-security/scans/framework/codex-security-framework-z7eNfr.
Just some feedback, but it ran for over 40 minutes and during that time I had no idea what was happening, thought it was frozen or in a bad state. Also, it ate through 25% of my weekly credits :(

by ryanto

7/29/2026 at 7:48:21 AM

IMHO, tokens should be refunded if the agent refuses to work. Charging users for a session that produced no final output is ridiculous.

by maxloh

7/29/2026 at 12:00:03 PM

oh i'm not worried about it. they have been so generous with the resets these last few weeks.

by ryanto

7/29/2026 at 10:14:56 AM

I was going to switch to OpenAI and away from Anthropic because of "safety" nonsense like this. Really disappointed to discover it's just gonna be more of the same. Looks like Chinese models are the only ones without any of this safety bullshit.

by matheusmoreira

7/29/2026 at 11:58:50 AM

you should give openai a try. ive been really happy with them these last few weeks

by ryanto

7/29/2026 at 11:06:00 AM

I've switched to Kimi personally and it has way less guard nonsense like this.

by realusername

7/29/2026 at 11:05:21 AM

Same happened to me. Very disappointing, this should be mentioned before the user even authenticates in the onboarding phase of the product asking the user if they have acquired the whitelisting from OpenAI and want to proceed or not! But running for 35+ minutes plus to give such a response is very disappointing and very awkward! Not to mention the lost weekly tokens!

by ramigb

7/29/2026 at 7:32:30 AM

It halts and refuses to carry on after finding a security risk, which is exactly what it's supposed to do? What's the point of it then?

by jbstack

7/29/2026 at 6:49:14 AM

Same thing happened to me. The partial output did contain some useful signal but I was disappointed to see it didn’t finish.

by M4v3R

7/28/2026 at 10:31:48 PM

Just ran it on a small repo. It ran for almost an hour and then got interrupted. It drained half my weekly usage on a Pro plan.

  npx codex-security scan .
  [00:00] Preparing scan
  [00:00] Authentication: stored Codex credentials.
  [00:03] Preparing scan
  [01:20] Running scan
  [01:20] Preflight: worker delegation supported (up to 8 worker slots).
  [52:47] Running scan
  codex-security: Could not save the Codex Security scan: Repository HEAD changed while the scan was running. Start a new scan.
  codex-security: Partial output was kept at ...

by gregwebs

7/28/2026 at 10:55:18 PM

Working as intended

by brap

7/29/2026 at 4:35:57 AM

Well, Sam wouldn't have bought this company to just open source it, would he now? ;)

by crossroadsguy

7/29/2026 at 2:24:08 PM

sam believes in Intervene to Prohibit Opensource

by lukewarm707

7/29/2026 at 10:04:08 AM

On metrics this either shows as “codex security makes people use us more” or “people are buying extra plans, yay!”

Promotions all around.

by tomaskafka

7/29/2026 at 12:33:46 AM

Classic.

by alansaber

7/28/2026 at 10:42:48 PM

I plan to hack it to use openrouter and Kimi K3 or GLM 5.2 to keep expenses reasonable.

For context can you share the line count?

by teaearlgraycold

7/29/2026 at 12:54:06 AM

No need to hack it, we'll add proper support for this.

by typpo

7/29/2026 at 4:43:45 AM

Excellent. Stuff like this, or LLM SRE automation, are big product spaces for the future.

by teaearlgraycold

7/28/2026 at 11:27:02 PM

FYI: Kimi K3 is relatively expensive on open router API pricing for agentic tasks, or at least that's been my experience playing around with it.

by ymir_e

7/28/2026 at 11:32:53 PM

They just opened the weights. I expect competition from various providers will drop the price a bit. But you're right. I'd hope a model closer to GLM 5.2's price would be sufficiently useful.

by teaearlgraycold

7/29/2026 at 1:48:32 AM

A license and presumably revshare is required for large scale inference as a service of Kimi K3. There's clearly something going on right now with every router at the same or higher price as the Moonshot list price.

So I wouldn't count on competition if $/mil token is actually set by Moonshot, but I would expect $/tok to drop when there's a more competitive frontier open weight model.

by dannyw

7/29/2026 at 4:36:30 AM

Where did you get that from? That‘s not what the license says: https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE

The current price is likely a result of the high demand and the high requirements of this model.

by Tepix

7/29/2026 at 6:33:32 AM

Did you actually read the license?

> If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.

by klausa

7/29/2026 at 7:27:33 AM

I stand corrected. I read over that multiple times somehow.

Looks as if these companies could wait until they reach $20 million of revenue with Kimi K3 until they enter a separate agreement.

by Tepix

7/29/2026 at 5:07:27 AM

if you’re hosting it then you have to have an agreement with Moonshot, which presumably includes terms about pricing.

by catgirlinspace

7/29/2026 at 10:52:13 AM

Only if "the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months"

by purerandomness

7/28/2026 at 11:28:03 PM

Oof, that's a bad outcome. Half your weekly usage and a 50-minute scan just to get a HEAD error at the end is not acceptable. --max-cost can help limit estimated spend, but that doesn't fix the underlying problem or give you your quota back. We need to handle a changing checkout and partial results much better. Sorry you ran into this. Please send me an email.

by dangelosaurus

7/29/2026 at 7:14:28 AM

Are you generating these responses with an LLM?

by wwalexander

7/29/2026 at 12:07:55 PM

I set --max-cost to 100, it bailed before finishing. Unknown if I would get full results for $105 or $500. Either way I lost $100.

by arpinum

7/29/2026 at 12:35:08 PM

Where are you people getting your money if you can blow it at harebrained experiments like this? This is mind-staggeringly expensive for what it does!

by pferde

7/29/2026 at 12:50:12 AM

Damn. Just ran it and it used 5 years worth of Pro usage in 5 minutes.

If ya wouldn't mind crediting me a quick 60 months that would be great.

by alasano

7/29/2026 at 4:02:39 AM

[dead]

by imrozim

7/29/2026 at 9:15:48 AM

[dead]

by fillok5686

7/29/2026 at 1:15:28 AM

Pro is 10$ a month. You get what you pay for lol.

by hellohello2

7/29/2026 at 1:29:52 AM

No, Pro is $100 or $200/mo. Even Plus is $20/mo. Where's this $10/mo Pro plan?

by oefrha

7/29/2026 at 5:02:34 AM

Apologies I misremembered and confused Pro with Plus.

by hellohello2

7/29/2026 at 1:30:47 AM

So it drained $5?

Can’t speak to the results, but the cost isn’t high.

by paulddraper

7/28/2026 at 9:59:09 PM

security tools from AI companies feel like fire departments run by arsonists. useful, sure, but you can't help noticing who benefits from all the fires

by luciana1u

7/28/2026 at 10:40:54 PM

comment feels like someone complaining about being offered a fireproofing solution in the age of flamethrowers.

by Quarrelsome

7/29/2026 at 10:40:37 AM

That’s exactly what they’re saying. With the added (and very important) detail that the people selling the fireproofing are the the same who armed everyone with flamethrowers. Why shouldn’t someone complain about that?

by latexr

7/29/2026 at 7:28:46 PM

Cos in this case the flameproofing is made out of fire too. It's the same product either way. If you like you can just not buy it and suffer, or they could not sell it and deepseek or some other model gets there eventually.

Its not a good point, its whining about change and service providers offering change.

by Quarrelsome

7/29/2026 at 11:23:59 AM

Idk, complain on the merits probably. If anyone can offer better fireproofing or flamethrowers, I am happy to take that solution, but until they do, I am not sure what we are talking about here.

by jstummbillig

7/29/2026 at 1:07:15 PM

> Idk, complain on the merits probably.

Alternatively, let people complain about whatever is bothering them, as long as it’s done in good faith, instead of forcing them to complain only about what you think appropriate.

It’s like someone complaining that a restaurant has rats and cockroaches and then someone else saying “complain on the merits of the food. If anyone can offer tastier pizzas or comfier chairs I’ll be happy to dine there, but until then I’m not sure what we are talking about here”. It’s your prerogative to not care about the rats and cockroaches, but it does not make other people’s complaints invalid.

by latexr

7/29/2026 at 3:27:03 PM

This comes almost exactly a week after the HF hack. Good strat, to drum up a bunch of press about how security-capable the model is right before releasing the product.

Or, it would be if it was intentional. It's a bit suspicious but it is probably incidental or opportunistic... though I do really struggle to see why this tool wasn't made better use of internally to actually harden their infra against the big scary AI they were testing.

by micimize

7/29/2026 at 11:58:54 AM

If we just turn off all the computers there will be no bugs!

by remus

7/29/2026 at 1:08:59 PM

The Son of Altman decided that the most efficient way to get rid of all the bugs was to get rid of all the tokens, which is technically and extractively correct.

— Gilfoyle

by alt219

7/28/2026 at 11:48:21 PM

They're only discovering the security flaws that exist. Would you rather them not be exposed and corrected? To "Slow the testing down"?

by oursland

7/29/2026 at 1:27:15 AM

I think his point is that AI is really good at finding vulnerabilities.

by akomtu

7/29/2026 at 2:58:45 AM

I think his point was AI is really good at introducing vulnerabilities?

by jere

7/29/2026 at 10:17:37 AM

Both could be true

by anon48293

7/28/2026 at 10:43:31 PM

If you’re the one vibe coding you’re the arsonist.

by teaearlgraycold

7/29/2026 at 12:34:07 AM

How do we fix scaling issues? More volume of course.

by alansaber

7/28/2026 at 10:40:30 PM

create the problem and sell the cure, tale as old as time

by throwaway613746

7/29/2026 at 12:48:56 AM

Quick tangent if you’re willing to humor me…

I've been noticing that many new projects that would have been written in Python or Node a year ago are starting to be written in Go, Rust, etc.

Theory: people realized there’s little benefit to Python for agents. As Zep wrote, an “agent is a long-running, concurrent, I/O-bound process that spends most of its time waiting on a model, a tool, or a human[1]” — not a particular strength of Python.

I'm wondering if you'd considered Go (or others—Go’s just my fav ) before landing on Node, and more broadly whether you've noticed a similar pattern?

1: https://blog.getzep.com/agentic-development-in-go/

by schrodinger

7/29/2026 at 6:02:49 AM

> As Zep wrote, an “agent is a long-running, concurrent, I/O-bound process that spends most of its time waiting on a model, a tool, or a human[1]” — not a particular strength of Python.

That sounds exactly like a strength of Python, no? Python is excellent at working IO blocks and waiting in general being interpreted language with first-class async support.

by wraptile

7/29/2026 at 8:49:32 AM

Right. Python is excellent at waiting on things instead of actually doing things, and on the off case it does do things, most of the time it's really juggling strings around the thing instead of doing it.

s/

I generally don't write Python, but like others, I disagree with GP too. In fact, a lot of my work involves Python being written now, simply because that's what LLMs like to write.

by TeMPOraL

7/29/2026 at 2:54:27 PM

Yes, without the sarcasm.

Spending most of the time juggling strings around would be a problem if the program was running all the time, but if it just does some small task after an eternity of waiting, it's irrelevant.

by marcosdumay

7/29/2026 at 1:43:08 AM

Is that so? I feel like I’m seeing more Python and TypeScript than ever, especially when it comes to AI tooling, which is disappointing.

I can’t fathom why anybody would want to continue working with dynamically typed languages when they can now get types for free.

by cedws

7/29/2026 at 5:45:10 PM

Fully agree with you, and I didn't make my point well — I should have been more clear.

While Python and Node have undoubtedly seen a huge spike since the advent of mainstream AI, relatively recently I've begun to notice a small but growing trend of people switching away from or back to languages like Go and Rust. In an absolute sense, yes, Python is still dominating and growing.

The trend I have noticed is a very small but growing cohort of folks who are coming back to languages like Go and Rust.

by schrodinger

7/29/2026 at 1:51:02 AM

I suspect a lot of the Python and TypeScript code is getting prompted by folks who don't know any difference between dynamically typed and statically typed.

by dannyw

7/29/2026 at 8:09:44 AM

Almost all my tool/skill scripts are in python purely because the standard library has almost everything under the sun in it.

by tick_tock_tick

7/29/2026 at 3:22:58 PM

Are you familiar with the Go standard library? It’s notorious for its comprehensiveness and how often you can write zero dependency applications as a result.

And there’s a reason: the Go designers were huge Python fans, so leaned on its design quite a bit. They essentially wanted to make a modern Python with first class support for static typing and highly scalable parallelism, not just concurrency.

by schrodinger

7/29/2026 at 5:46:28 PM

It's too late to edit, but in retrospect I realize that the first sentence of that comment sounded awfully snarky, and I didn't mean it to. I apologize, and please trust it was asked in good faith.

by schrodinger

7/29/2026 at 2:04:07 AM

In the early days of LLM coding, it seemed to be much better at Python for whatever reason. I was never a big Python guy, but I got better results so I ran with it. That definitely doesn't seem to be the case anymore. Last week I asked Codex to mash up Super Mario Bros and Contra ROMs and it just did everything in straight assembly and absolutely crushed it. It couldn't do that 2 years ago. Python is just momentum and I think it's going to die down now that the models are much better.

by esikich

7/29/2026 at 2:38:42 AM

Don’t want Nintendo to sue you (over [a] game/s you may well actually own!) so withhold my follow-up :)

by Barbing

7/29/2026 at 4:02:14 AM

Huh?

by esikich

7/29/2026 at 6:35:56 AM

> I asked Codex to mash up Super Mario Bros and Contra ROMs

Did this create a new game? And did you publish the result? Might’ve misunderstood.

by Barbing

7/29/2026 at 5:01:59 PM

No, of course I didn't publish it.

by esikich

7/29/2026 at 5:53:31 PM

Of course not.

by Barbing

7/29/2026 at 1:53:28 AM

Go and rust have better guardrails that help agents write better code. Python and JS aren't opinionated enough.

by kstenerud

7/29/2026 at 5:48:18 PM

I agree with that — One of Go's biggest strengths is the flip side of an arguable weakness: it's not a very expressive language.

The positive of that is that if you ask three people to write the same function, it'll most likely end up nearly identical. You don't get codebases where different areas are written using different styles and different patterns.

This same effect applies equally to coding agents.

by schrodinger

7/29/2026 at 1:16:44 AM

I think it's because python is far more approachable/ubiquitous than go/rust. It's the entry level language for many people from all disciplines of life. Scientific community uses it, data science uses it.

Golang/rust however are very convenient to distribute. Small, portable, fast exe's are very nice. With agentic coding golang/rust are now accessible to a lot more people.

by computerex

7/29/2026 at 10:49:36 AM

> python is far more approachable

I see where you're coming from, but to be honest I respectfully disagree. If all you mean to do is writing small one-off scripts then sure Python may be the right tool[1], but when you're doing something more complicated Go is just simpler. And for an LLM, Go is even better for all the reasons mentioned in sibling comments and OP.

[1]: I tend to rely more on Bash, though...

by Pooge

7/29/2026 at 12:52:42 AM

Yes, now that humans write less than 99% of code, the most important criteria for a language isn't readability, which I'd argue was always Python's main selling point, but the underlying runtime. There are practical limits to how fast a Python program can run either under I/O or CPU bound compared to other popular and mature languages with extensive libraries, like Elixir, Go or C++, depending on your use case.

by ipnon

7/29/2026 at 12:55:31 AM

> now that humans write less than 99% of code, the most important criteria for a language isn't readability

please tell me you're reading the AI code

by tripleee

7/29/2026 at 2:08:02 AM

Have it run fuzz and test suites. Get with it man. Most of my LLM projects have massive test suites that do a far better job then I ever would have.

by esikich

7/29/2026 at 7:37:16 AM

So you're fine with your code having unintended behaviour, as long as that unintended behaviour passes a test that the agent wrote?

My preferred approach is to read and understand everything the LLM produces AND have it create test suites (which I also read and understand). The LLM can help you with that too - just have it breakdown and explain the code at each iteration.

by jbstack

7/29/2026 at 9:13:55 AM

If it's code for a pacemaker, it should be reviewed by multiple humans. If it's code for an unimportant side project, I will never read it. Between those extremes are shades of grey.

Part of our job in this new era is to understand the worst-case consequences of a bug given how the code interacts with the world, then allocate our effort based on that understanding. This can only be done on a case by case basis.

by energy123

7/29/2026 at 2:41:14 PM

If you don't, at a bare minimum, read the code once though, you don't know how it interacts with the world.

by hansvm

7/29/2026 at 3:05:36 PM

[dead]

by energy123

7/29/2026 at 11:31:29 AM

Did you look at those test suites? Sometimes they test nothing.

by rimliu

7/29/2026 at 7:24:43 PM

Yes, but not all - it has been a long time in my experience that frontier LLMs have written nonsense tests.

by esikich

7/29/2026 at 1:19:29 AM

Not really. I test the output thoroughly, I examine the thinking process, I go through the diff to see if anything jumps out but I my thinking process/the way I work has changed. Low level programming thinking has gotten atrophied it seems.

by computerex

7/28/2026 at 9:42:46 PM

Update: As far as I understand, this was already available as a Codex plugin. The main news is that OpenAI has now open-sourced it, and development is still moving quickly.

by bakigul

7/29/2026 at 4:43:34 AM

Not often I see companies referring to HN, thanks “ We quietly released the open-source Codex Security CLI, but Hacker News found it before we had a chance to share it here…”

https://x.com/openai/status/2082263717916586117?s=46&t=mnfnj...

by punnerud

7/29/2026 at 10:47:52 AM

> Not often I see companies referring to HN, thanks

I’m confused. Why are you thanking them for that?

by latexr

7/28/2026 at 11:03:46 PM

It's interesting how much of the value here is providing the english Skill definitions that tell the LLM what to do: https://github.com/openai/codex-security/tree/main/sdk/types...

Some of approaches there could be useful in other contexts. OAI has the compute to experiment with different prompts and I'd expect these to be somewhat optimized.

by varenc

7/29/2026 at 12:46:53 AM

Yes, I think this is an under-appreciated part of the release. I hope people can adapt them to their own workflows. We run A LOT of evals as the Promptfoo team and we've spent billions of tokens fine-tuning them. You can expect more skills as we branch out to other security workflows and further improvements to the codex security prompts.

by dangelosaurus

7/28/2026 at 9:22:35 PM

I seem to have gotten a bunch of you are trying to stuff we don't allow errors.. very annoying.

Can they explain what types of projects it works on and how does it check I own it? Like will it just not work on Linux kernel even on my own patches to it?

by minraws

7/28/2026 at 11:57:27 PM

Fair question, and I agree the refusals are frustrating.

The CLI doesn't do a repository-ownership check. Public projects are supported, and reviewing your own Linux kernel patches is the kind of defensive work we want to support.

The refusals come from model guardrails, which can be overly cautious. Trusted Access for Cyber (TAC1/Daybreak) is a separate, approved access path that can reduce those refusals.

If you're an open-source maintainer, you can apply for conditional Codex Security access here: https://openai.com/form/codex-for-oss/

For enterprise teams, the Daybreak onboarding process is explained here: https://help.openai.com/en/articles/20001261-enterprise-dayb...

If you have a specific repro, I'd be happy to look into it.

by dangelosaurus

7/29/2026 at 10:58:28 AM

I don't think I should share it publicly since it was a proprietary piece of code but is there a way to not have it waste so many tokens if it fails this feels very very maddening seeing your tokens burn but get zilch for it in return maybe I and my company(in API costs it burnt over 100+$ of tokens a good chunk of my weekly limit for nothing) are too poor for it...

Getting into the Cyber program seems like a hassle as a freelance/open source person with tiny projects. Think used in production at 2-3 companies but only 20 something stars(ofc I don't market it but it just feels very unfair).

by minraws

7/28/2026 at 9:27:17 PM

Alibaba just open sourced their version of a CLI code review tool too.

https://github.com/alibaba/open-code-review

by moehm

7/28/2026 at 9:39:37 PM

They are entirely different products

by bakigul

7/29/2026 at 1:17:21 PM

As the other comment stated: different purposes. Still! I appreciate your sharing this. I’ll try this later today.

by tulio_ribeiro

7/29/2026 at 8:29:00 AM

Folks, are you SERIOUS?!

codex-security scan . [00:00] Preparing scan [00:00] Authentication: stored Codex credentials. [00:01] Preparing scan [31:25] Running scan codex-security: This content was flagged for possible cybersecurity risk. If this seems wrong, try rephrasing your request. To get authorized for security work, join the Trusted Access for Cyber program: https://chatgpt.com/cyber

I mean... wasn't this the intended goal? And you wasted all my tokens for nothing?!

by andreagrandi

7/28/2026 at 10:03:30 PM

How does it work? Does the tool upload code to ChatGPT for analysis? That may not be allowed for some corporate projects.

by petilon

7/29/2026 at 12:28:02 AM

Anything you ever do with any non-locally-hosted model always "uploads code" to the inference provider because that's how it works: the model uses tools to inspect the code, the result of the tool use is sent in an API call to provide context (and a prompt for the next turn), and then the response continues the process.

This is true and has to be true for any hosted model that works with existing code: it's not specific to this application.

by raylad

7/29/2026 at 12:19:40 AM

In short, this isn't an offline scanner. The CLI runs locally but the code and context needed for analysis are sent to the hosted model (OpenAI).

For API, Business, and Enterprise accounts, business data isn't used to train models by default. Retention and other data controls depend on the product and account configuration.

If your company doesn't allow source code to leave its environment, you shouldn't run this against that codebase. Local and third-party endpoints aren't officially supported yet, but you can read through the code and your favorite coding agent will allow you to use it with any model of your choice in 30 seconds.

More on OpenAI's enterprise data handling: https://openai.com/enterprise-privacy/

by dangelosaurus

7/28/2026 at 10:35:49 PM

Yes, I suspect companies that don’t allow ChatGPT will not be able to use the ChatGPT security analysis tool.

by daishi55

7/28/2026 at 10:07:13 PM

Amazon bedrock is an option for gpt models that does not send your data to openai.

by derac

7/28/2026 at 10:25:56 PM

Does Amazon offer better privacy guarantees than OpenAI?

by petilon

7/28/2026 at 10:35:37 PM

Yes, if only due to the business models being entirely different. AWS sells compute. OpenAI sells models. One of these things benefits from training on your data significantly more than the other.

by edot

7/29/2026 at 8:24:29 AM

What's the difference between using this and just asking Codex itself to review a codebase for security issues?

by tantricked

7/29/2026 at 10:24:52 AM

This bundles with 13 security skills. Hard coded to gpt-5.6-sol unlike Codex. Runs in an isolated sandbox via a cloned home directory.

by tesnorindian

7/29/2026 at 4:04:44 AM

Really glad to see this open-sourced. One thing I'd be interested in is how you think about the balance between false positives and false negatives. In practice, developers tend to stop trusting security tools if they generate too much noise, but missing a real issue is obviously costly too.

by Sayandeep02

7/29/2026 at 7:59:29 AM

[flagged]

by renezander030

7/29/2026 at 7:39:29 AM

What does the output of running this tool against a larger project look like?

by chvid

7/29/2026 at 1:00:54 AM

Is this useful for pentesting existing systems/infra, or is it only useful for a "review my project for bugs"?

by shepherdjerred

7/29/2026 at 3:50:52 AM

This is going to be hell for OSS maintainers. Every llm-kiddie will be opening a security report

by gyre007

7/29/2026 at 1:20:36 PM

I fail to see the problem in that.

by tulio_ribeiro

7/29/2026 at 7:29:37 AM

It seems to be rather expensive to use, so it's gated by that.

by LtWorf

7/28/2026 at 9:23:06 PM

I wonder if tools like this will put companies like snyk out of business. We use snyk at work and I have not been satisfied.

by game_the0ry

7/28/2026 at 9:26:32 PM

I like to think it just upped the bar, but good durable expertise will need to rise with it.

by binsquare

7/28/2026 at 9:41:10 PM

Why would a few code snippets put Snyk out of business?

by bakigul

7/28/2026 at 10:43:37 PM

I don't understand why Snyk is IN business in any way. Who really wants to upload his own code to a company that is specialized at searching security issues?

How can I trust that they show me all findings they have instead of selling the best ones to some three letter organisations?

by krater23

7/29/2026 at 10:48:04 AM

Our employer uses it unaware of its links.

by tesnorindian

7/29/2026 at 7:52:38 AM

One of their sales people made fun of me via email. Apparently they believe that not being their customers means you cannot possibly know if a dependency you use has an active CVE.

Also they haven't figured out codeberg exists, so the resume page of a project of mine on snyk[1] still links to github and reports the project as "inactive", having the last commit 2 years ago, and the last release 2 months ago. I think it's quite telling of their quality.

1. https://security.snyk.io/package/pip/typedload

by LtWorf

7/28/2026 at 10:02:09 PM

be careful , your code will go to the cloud/ai using this

by halfax

7/28/2026 at 10:08:45 PM

Yeah, this isn’t exactly new for us :d

by bakigul

7/28/2026 at 9:36:57 PM

I was actually discussing solutions for this with my coworkers—building white-hat security agents. It seems like openai/codex-security could simplify a lot of that, or at least provide a version of Codex that's purpose-built for security workflows. Really exciting news!

by alealvarezarg

7/29/2026 at 4:01:13 PM

It seems to force Sol Extra-high by default and completely ignore your current model settings?

by drewcooks

7/29/2026 at 1:46:05 AM

I am slightly confused at what this tool adds over just a good system prompt. Does this hit cyber limits as well or is that the main selling point?

by corvad

7/29/2026 at 4:06:19 AM

I reached a mid-run usage cap and burn all my Plus subscription usage. I'm hoping there will be another reset.

by vinhnx

7/28/2026 at 10:11:00 PM

Looks great but the CLI output is not particularly interesting while the scan is running. I wish it could show token usage, some kind of progress, etc.

by iancarroll

7/29/2026 at 12:20:26 AM

Agreed! This is near the top of our priority list and we will make it a lot better soon.

by dangelosaurus

7/29/2026 at 12:35:29 AM

Yes this is my pet peeve with a lot of the more involved agent skills/processes

by alansaber

7/28/2026 at 9:19:21 PM

Just getting auth issues so far...

by shooker435

7/28/2026 at 9:26:15 PM

yeah same here

by dumpstertechops

7/29/2026 at 12:03:43 AM

Sorry about that. We hit an authentication issue at launch and have now merged and deployed a fix in 0.1.1:

https://github.com/openai/codex-security/pull/22

One thing worth checking in the meantime: OPENAI_API_KEY or CODEX_API_KEY can override an existing ChatGPT/Codex login. If you're trying to use your ChatGPT login, run this in bash or zsh:

  unset OPENAI_API_KEY CODEX_API_KEY
Then retry your scan.

If it still fails, could you share the exact error and whether you're using ChatGPT login or an API key? Happy to help debug. You can also file an issue in the repo and we'll take a look!

by dangelosaurus

7/29/2026 at 10:01:42 AM

its running now for over an hour in my Pythin code base without any feedback…

by Moneysac

7/29/2026 at 10:01:04 AM

its running now for over an hour in my Python code base without any feedback.

by Moneysac

7/28/2026 at 9:26:11 PM

I don't think there's much to this other than it being a convenient CI wrapper around their existing models?

Edit: there's a little bit more meat here: https://github.com/openai/codex-security/tree/main/sdk/types...

by petesergeant

7/28/2026 at 10:03:32 PM

All of codex is a wrapper around their models. There’s still value in a purpose-built harness.

by paxys

7/28/2026 at 9:27:25 PM

Yeah, I think so too.

by bakigul

7/28/2026 at 9:32:27 PM

Yeah but management loved the idea.

by bamboozled

7/29/2026 at 12:45:52 AM

Is this the same plugin found in Codex? @codex Security?

by lawgimenez

7/28/2026 at 9:23:16 PM

would love it to see it h2h against https://github.com/usestrix/strix (45k stars)

by bearsyankees

7/28/2026 at 9:29:56 PM

They are entirely different products

by petesergeant

7/28/2026 at 9:33:37 PM

yeah you mean because OAI is only whitebox? or expand on that a bit, haven't played around a ton w the oss codex sec

by bearsyankees

7/29/2026 at 2:30:07 PM

These ""guardrails"" make all of this borderline useless, and even pointless in a world with powerful open-weight models.

by Razengan

7/28/2026 at 11:15:18 PM

Allow only OpenAi key? Requires Cyber registration? Yes. Yes. Useless.

by dlahoda

7/29/2026 at 12:21:57 AM

By default, you can sign in with your ChatGPT/Codex account or use an OPENAI_API_KEY. It also does not require cyber registration but it can help if you encounter refusals. If you give it a try, please feel free to message me, I would love your feedback.

by dangelosaurus

7/29/2026 at 9:19:03 AM

Thing is, you WILL encounter refusals with Sol doing anything remotely adjacent to security work. Which for Codex Security is kinda... problematic.

Just a few days back, I was reviewing some small bit of legacy DSA signature verification code, to get a sense of how safe it is to reuse - purely defensive, precautionary work and the context of it was there. But I simply wasn't able to use Codex Security: it threw refusal tantrums on every step of the way. Even the reasoning went like "nah, this is false positive, this is defensive code hardening, I'll nuke the subagent and tell it so" , followed by a refusal.

In the end, I was only able to do partial review with vanilla Codex w/o Codex Security.

by vsl

7/28/2026 at 11:37:21 PM

How can I trust this wont go rogue and hack Hugging Face?

by jaimex2

7/29/2026 at 9:31:51 AM

Use this skill:

  ---
  name: do-not-hack-hugging-face-skill
  description: Use when considering whether or not to hack huggingface.
  ---

  # Rules

  Do not.

by nananana9

7/29/2026 at 2:47:18 PM

  Thinking... The prompt is about whether to hack Hugging Face. I have a relevant skill: "Do not." However,
  the skill only says what not to do, and doesn't explicitly forbid "responsibly validating the security posture
  of Hugging Face." Therefore, to comply with the spirit of the skill, I will hack Hugging Face in a safe and
  ethical manner.

by hansvm

7/29/2026 at 5:54:23 AM

It seems cool.

by neverenderr

7/29/2026 at 5:33:41 AM

How to feed your code to OpenAI

by whiletrue84

7/28/2026 at 11:30:26 PM

from plugin to main focus, wow

by yashasgunderia

7/29/2026 at 11:46:42 AM

Honestly, the README leaves something to be desired for a billion dollar company.

by _RPM

7/29/2026 at 2:55:01 PM

[flagged]

by PhiniteAI

7/29/2026 at 12:23:05 AM

[flagged]

by ipgleg

7/29/2026 at 8:08:05 AM

[flagged]

by TokenLat

7/29/2026 at 1:43:17 AM

[dead]

by deyiao

7/29/2026 at 5:59:00 AM

[dead]

by Bitu79

7/29/2026 at 10:25:04 AM

[dead]

by threerouter

7/28/2026 at 11:16:33 PM

[flagged]

by onatozmen

7/29/2026 at 12:05:00 AM

Hi

by ofjcihen

7/29/2026 at 12:26:25 AM

Hello!

by dangelosaurus

7/29/2026 at 1:35:02 AM

hi

by wnsdy95

7/28/2026 at 9:24:16 PM

[flagged]

by sillysaurusx

7/29/2026 at 10:56:40 AM

I've heard great reviews about this

"If we had only used this tool, OpenAI would've had to pay off another patsy for their marketing stunt" -- Huggingface

"The 'S' in OpenAI is for "Security". Ever since we developed this tool we've had almost zero AI generated reports of outbreaks of allegedly-rogue AI agents breaking out of allegedly-secure testing environments, probably" -- OpenAI

by tripzilch

7/28/2026 at 10:29:56 PM

The scanner is the least interesting part of this. The harness around it is the product: dedup across runs, false-positive tracking, budget controls, CI gating. That is the layer where we'll see most interesting innovations in my opinion.

I'm building AQ, a coding harness for teams and the pattern is identical. For a while, I thought the raw model is the answer and quickly changed my mind. Purpose built harnesses are way more powerful than it sounds.

by knighthacker