Mythos Finds a Curl Vulnerability

5/11/2026 at 7:33:34 AM

Quote:

"My personal conclusion can however not end up with anything else than that the big hype around this model so far was primarily marketing. I see no evidence that this setup finds issues to any particular higher or more advanced degree than the other tools have done before Mythos. Maybe this model is a little bit better, but even if it is, it is not better to a degree that seems to make a significant dent in code analyzing."

It's a good reminder for us all that the competition in this space is rough and lots of more or less subtle marketing is involved.

by rzmmm

5/11/2026 at 9:34:01 AM

Anthropic using marketing to convince people their models are more advanced, better built, or that AI is a threat that needs to be regulated because only they have the answer? I’m shocked.

More seriously, so far I haven’t seen much indication that Mythos is more than Opus with a security focused code analysis harness. That said, the fact it can find these bugs in an automated fashion is the more important takeaway outside of the hype.

I’m curious what the error rate is on the detections, because none of that means much if it is wrong 90% of the time and we are only hearing about the examples that are useful marketing.

by therealpygon

5/11/2026 at 9:59:22 AM

>> Anthropic using marketing to convince people their models are more advanced, better built, or that AI is a threat that needs to be regulated because only they have the answer? I’m shocked.

I remember when OpenAI was saying GPT-2 was too dangerous to release.

by johnbarron

5/11/2026 at 10:44:41 AM

I remember when there was a guy at Google years a few years ago that was convinced that they had an internal, sentient creature in their labs (I think maybe 4 years ago?)

If I’m not mistaken, after the media cycle, he lost his job for breaking confidentiality.

That was the opposite of marketing, Google really didn’t get how to turn this into a product until ChatGPT happened.

by stingraycharles

5/11/2026 at 1:00:05 PM

They most likely understood that it wasn't viable for anything. OpenAI just yolo'd it and now we're dealing with the fallout. I'm fairly certain that any management layer at google isn't going to say yes to "invest 5 billion to make 10 million" scheme that OpenAI, Anthropic, are currently running.

by jabwd

5/11/2026 at 2:22:17 PM

I for one cant wait for the 10 million to go all the way to zero

by gofreddygo

5/11/2026 at 5:52:34 PM

"ChatGPT has over 900 million weekly active users worldwide. ... ChatGPT Plus has around 50 million paying subscribers"

by mistrial9

5/11/2026 at 7:55:20 PM

What you have typed does not address anything the person you are responding to said.

With those 50 million subscribers, how much do they pay and how much do they cost? That is the only relevant piece of information when discussing the investment and returns of OpenAI.

by muwtyhg

5/11/2026 at 9:03:51 PM

> "invest 5 billion to make 10 million"

business is contextual, and is a game of numbers? If you agree, then there is a difference between "I made money selling lemon drinks at my driveway, but I sold a car to make room" .. versus "I have recurring revenue of 50 million x $80 USD per month, and it is growing, and I am using cheap credit to build that" .. Numbers have a meaning, and the larger dollar recurring revenue cannot be matched in any way, no matter how much I spend. IIR ChatGPT is the fastest adopted software in the history of the Internet.

by mistrial9

5/11/2026 at 11:25:02 PM

Is it growing?

Don't they report annualized revenue AKA the best month times 12? How is that comparable?

by JetSpiegel

5/11/2026 at 10:28:32 PM

They have no moat.

by jabwd

5/11/2026 at 2:25:27 PM

Google is the leader, they really don't want AI to be a success, it only comes with a risk of disruption. They probably don't even really believe it's going to be that big of a deal. They are only in that game to hedge; sure they have wasted a trillion dollars if AI doesn't come through, but they will earn that back in 3-5 years. So why would they need to do deranged marketing stunts and sacrifice their credibility for that?

If OpenAI or Anthropic doesn't turn this into a trillion dollar industry FAST, they are cooked. The strategy of building up fear around your product is risky, but necessary. There is simply no way to grow the AI business fast enough if they can't talk directly to the CEOs and bypass input from the employees, and baba yaga stories are perfect for that. Every time the CEO hears an employee say that the AI isn't working great for him, he hears an employee that's scared for his job or for his life, dismisses it, and sends out a mandate that everyone needs to prompt an AI every time they as much as need to go to the toilet.

by MadxX79

5/11/2026 at 11:17:19 AM

[dead]

by player1234

5/11/2026 at 3:03:52 PM

Context from 2019: https://en.wikipedia.org/wiki/GPT-2

>While previous OpenAI models had been made immediately available to the public, OpenAI initially refused to make a public release of GPT-2's source code when announcing it in February, citing the risk of malicious use;[8][5] limited access to the model (i.e. an interface that allowed input and provided output, not the source code itself) was allowed for selected press outlets on announcement.[8] One commonly-cited justification was that, since generated text was usually completely novel, it could be used by spammers to evade automated filters; OpenAI demonstrated a version of GPT-2 fine-tuned to "generate infinite positive – or negative – reviews of products".[8]

>Another justification was that GPT-2 could be used to generate text that was obscene or racist. Researchers such as Jeremy Howard warned of "the technology to totally fill Twitter, email, and the web up with reasonable-sounding, context-appropriate prose, which would drown out all other speech and be impossible to filter".[18] ...

by neuronexmachina

5/11/2026 at 5:19:32 PM

It's kind of funny watching the behavior on the forum of different groups with different beliefs.

"AI can't do anything harmful at all, kick this shit up to 11. It's all marketing, bla bla"

and

"My grandma gave away all her money to AI bots and is now starving in the street. My uncle murdered his wife and is trying to get married to GPT-4o. He thinks they are going to elope to a data center on a tropical island and live happily ever after".

I think the 'AI can do no harm, it's marketing" people are really disconnected from reality and that any other product that behaved in the same manner would have been banned in most places.

by pixl97

5/11/2026 at 6:38:05 PM

AI chatbots have caused real harm. It has tragically convinced and encouraged a number of people to commit suicide, to say nothing about scams. It is having a real effect on the social fabric of our society.

I don't understand what point the people who blame the dangers of AI on marketing.

by abustamam

5/12/2026 at 4:55:32 AM

The sociocultural dangers weren't the danger they were referring too, Claude Mythos was purported to be so powerful that if released to the public it would result in all software being 0-dayed and so they could only give select important groups access. Curl's analysis said ehh, it didn't really seem that much better.

Now people who are getting negatively affected because they think AI is more real and more intelligent than it actually is and get tricked by it, well that is dangerous but for different reasons.

by robotbikes

5/11/2026 at 8:57:30 PM

> I remember when OpenAI was saying GPT-2 was too dangerous to release.

The world didn’t end yet - but did it improve?

by DANmode

5/11/2026 at 10:37:11 AM

"it can almost like write 2 paragraphs!" "It might be conscious" "this is basically AGI, we had to fire someone who spilled the beans"

by 2ndorderthought

5/11/2026 at 11:33:27 AM

I always thought he was fired for making crackpot statements to the press in reference to his professional capacity, and thus creating bad PR and embarrassing spectacle for his employer. Seems like legitimate reasons to me.

by etiam

5/11/2026 at 11:38:26 AM

An interesting question now is whether he had standard mental health issues, or if he was an early example of AI psychosis or whatever we call people who are falling in love with their AI chatbots because they tell them how smart they are.

by ZeroGravitas

5/11/2026 at 12:11:43 PM

Considering Richard Dawkins has recently succumbed to the same delusion it is a reminder that no matter how intelligent someone may otherwise be, we are all human and have certain tendencies and blind spots; anthropomorphizing non-entities being one of those.

by paradox242

5/11/2026 at 12:18:24 PM

Richard Dawkins is 85 to be fair, just like Bernie Sanders is 84 when he made similar comments.

The other guy worked on Google's AI safety team where one would expect he'd have a basic grasp of how the technology works before making outlandish claims.

by dmix

5/11/2026 at 1:10:17 PM

One phenomenon that spooks me is when intelligent people believe in idiotic things.

It makes me wonder if there's a wrong turn in the road that I too might fall in the same pit.

by surgical_fire

5/11/2026 at 1:54:39 PM

Vigilance is warranted, I think.

I can't find it right now, but something came up a few years ago (probably on HN) about highly intelligent people being more adept at making up arguments to rationalize beliefs and actions that they had taken for other reasons entirely.

Sort of makes sense that wielding a more complex mind would offer more complex ways to go wrong, doesn't it?

by etiam

5/11/2026 at 2:23:32 PM

And on balance, it also can mean that they make connections and see truth where others only see the facade. Both statements can (and are true) because highly intelligent people are still just people. Some people’s “delusions” are absolutely correct, and others “facts” are nothing more than anecdotes told to convince themselves of what they want to believe.

Sounds more like “intelligence” isn’t the only defining metric for such behavior to occur in people, because that describes a lot of less intelligent people too. Though, I suspect highly intelligent people are at least somewhat more likely to end up on the “correct” side of the facts.

by therealpygon

5/11/2026 at 1:56:37 PM

As someone who watched one of their heros fall for some stupid cult like thing ten years ago and wondered the same thing. Then many years later fell for some dumb stuff. The answer is you probably will. Try to stay intellectually flexible, it'll be okay.

by 2ndorderthought

5/11/2026 at 2:42:44 PM

I am afraid of that, I wasn't joking.

I have seen people I consider as much smarter than me fall for some very idiotic things. I certainly don't consider myself immune.

I think that the advice to try being intellectually flexible is a good one. Strive to learn new things, expose yourself earnestly to ideas that challenge your beliefs, exercise empathy, etc

by surgical_fire

5/11/2026 at 2:03:53 PM

Good point.

Optimization on "Human Feedback", early exposure to high-effort experimental systems... I wouldn't be surprised it that turns into a bigger field than is generally recognized today.

Looking at it from the outside, I think it's still pretty hard to see how he came to end up in that position, but with a bit of individual vulnerability, arbitrary time to boil the frog slowly, and a fairly large number people exposed, maybe it would be stranger not to have the event occur with someone.

by etiam

5/11/2026 at 4:52:45 PM

And Anthropic was founded by former, high ranking OpenAI employees so they were accustomed to the classic "its so dangerous we can't release it" trope.

It sounds like Mythos is good but none of us know exactly how good since they haven't released it yet. It also sounds like Anthropic is compute starved which is probably the biggest reason it has had a public release

by slipnslider

5/11/2026 at 12:10:20 PM

This is roughly what I was assuming but of course the big caveat here is that they were already using the existing LLM driven tooling on an extensively audited codebase.

So while anthropic's marketing may be hype there just wasn't much left to find, a point he makes in the blog post.

Whether it's a big step forward for other kinds of projects is difficult to tell, but this highlights that everybody should be using AI code review tools to audit their existing code today, and not everybody is.

by JeremyNT

5/11/2026 at 12:44:08 PM

None of those other LLM tooling made the claims they're too dangerous to be released and used though, unlike Anthropic did with Mythos.

What it highlights, is that Mythos doesn't seem so much better than other LLM driven tooling at finding security issues, which was the strongest claim Anthropic made in the first place.

by embedding-shape

5/11/2026 at 1:20:13 PM

People love defending Anthropics shortcomings…

“Mythos isn’t supposed to be that good at security, because actually Anthropic was referring more about running llms than mythos specifically”

“The opus model is worse because they have no compute because they are training mythos. The degraded performance is justified!”

“All the bugs in Claude code is just because the models are so good they are just looping and are shipping fast”

Constantly see people crawl out of the woodwork to defend a trillion dollars company overhyping every press release it gives

by iterateoften

5/11/2026 at 6:17:39 PM

If policitians can buy online supporters to manipulate perception during their election campaigns, I'd expect private corporations would too. Of course people can become very biased on their own, but an online PR/Marketing/Influencing campaign might encourage them to be more vocal.

by ASalazarMX

5/11/2026 at 7:59:28 PM

It's silly to act like they've got mud on their face when Mythos and Opus are apparently some of the very best models. Anyone that has found value out of previous LLMs is likely to find more value out of the newest ones. The only thing Mythos looks bad against is the very tall bar some people have imagined. People are putting too much weight on marketing and then reaction to marketing.

by AgentME

5/11/2026 at 8:04:11 PM

> People are putting too much weight on marketing and then reaction to marketing.

No, what others are doing, which I've done myself in the past too, is to evaluate how much their marketing matches up with reality, then share our experience about that. Very different than just "putting too much weight on marketing".

by embedding-shape

5/11/2026 at 7:52:27 PM

It's important to keep in mind that very, very few projects are as rigorously tested as curl, so while it's interesting to hear this feedback I think curl would be a torture test for any security scanning. I'd be more interested to hear about other random libraries that aren't as thoroughly analyzed as curl; show me some results for GnuTLS, for example, or dpkg/rpm/apt/dnf/pacman/etc.

by danudey

5/11/2026 at 9:40:14 PM

I think one of the points of TFA was that other AI tools found many vulnerabilities; after having fixed those, mythos did find another vulnerability the others missed, but that seems to imply this model is only marginally better than the competition instead of being on a different league altogether like it's marketed. Paraphrasing the author: sure mythos will find lots of security issues in gnutls, but so will gpt or opus (they acknowledge explicitly that all those tools are getting very good).

by p91paul

5/11/2026 at 5:00:19 PM

> None of those other LLM tooling made the claims they're too dangerous to be released and used though, unlike Anthropic did with Mythos.

I do think they've said similar things in the past, but regardless Anthropic's BS marketing is something to behold and viewing it with extreme skepticism is smart.

> What it highlights, is that Mythos doesn't seem so much better than other LLM driven tooling at finding security issues, which was the strongest claim Anthropic made in the first place.

That's the conclusion Daniel makes and it definitely seems plausible, his opinion absolutely carries a lot of weight with me for sure.

But I hedge a little because we don't really know how much human labor was required to supplement those earlier LLM-assisted reviews of curl, nor do we know how easy it was for the person who used Mythos to generate the new batch. So the kind of bug hunting that might be "possible but still labor intensive" via current tooling might be far easier to accomplish with less skilled developers using Mythos.

And who knows, maybe Mythos is better on worse codebases, curl benefits from being very good to start from :)

by JeremyNT

5/11/2026 at 12:55:17 PM

Actually, OpenAI made a similar claim about one of their GPT models a while ago…

Funnily enough that was while Dario Amodei was their research director.

by hug

5/11/2026 at 3:01:56 PM

If you're referring to gpt-2 in 2019, that primarily about concerns with it being used by spammers and fake content generators. In retrospect, that was a totally valid concern.

by neuronexmachina

5/11/2026 at 6:47:53 PM

They had a reddit with GPT2 back and forth I have to say I got suckered into a conversation before I figured it out -- it was definitely the OG Moltbook of non sequiturs

by jimmySixDOF

5/11/2026 at 8:18:17 PM

Too dangerous to be released, right after the Department of Defense* dropped them

by stirfish

5/11/2026 at 4:28:51 PM

Everyone should be using exclusively a proof assistant (Lean/Agda/Rocq/Isabelle) and proving their code correct, but they're not.

Do you see how ridiculous the zealotry sounds when its not your personal kind of zealotry?

by voxl

5/11/2026 at 10:08:55 AM

Curl simply isn't a good data point. It's one of the most picked-over codebases in existence with extensive security testing practices. All the researchers using not-quite-Mythos models have had plenty of time to report bugs up to this point. Daniel may be right that Mythos hasn't been a game changer for curl but the preconditions are different for virtually any other codebase. Perhaps the real marketing here is his own modesty about curl's maturity.

by thombles

5/11/2026 at 10:40:34 AM

To me, it is a very good data point.

Curl uses all sorts of tools, including AI tools to find bugs. These tools, according to the article found hundreds of bugs including a dozen CVE.

Mythos found one vulnerability. It means the Mythos is just another tool, not the revolution it claims to be.

It is common that when a new tool is introduced that a bunch of bugs are found, with diminishing returns. Mythos finding one vulnerability is consistent to what I would expect for a major update to an existing tool, which Mythos is over existing LLM-based solutions.

by GuB-42

5/11/2026 at 2:02:47 PM

I had a totally different take. The fact that Mythos found only one vulnerability is testament to how solid curl is, not how bad Mythos is.

Look at the Firefox blog post where they found something like 400 (or more) findings.

I have no doubt Mythos is very good at this, but I also don't think it's something unattainable by other labs within the next few months, with focus.

by atonse

5/11/2026 at 4:06:49 PM

The point is that Anthropic claims it’s a huge leap over everything else. But it isn’t.

by skywhopper

5/11/2026 at 4:26:24 PM

This depends on the actual number of undiscovered bugs still in curl. If there is nothing to find then even a 10x better Mythos will find nothing. Also I think the quality of the codebase matters a lot when it comes to finding bugs. Its possible that the curl is so well written that it is relatively straightforward for existing ai tools to find bugs.

by rohit89

5/11/2026 at 6:05:30 PM

But both things can be true. It could be a huge leap (see Firefox’s example) but also find almost nothing in an already well maintained and audited codebase, and that could mean there isn’t much to find.

by atonse

5/11/2026 at 7:23:19 PM

Okay, but how do we know that all 400 plus hits were actual vulnerabilities? I didn't read too deeply into it so I might've missed something but did someone test and validate each of those vulns to confirm that they were actually vulns?

by ethin

5/12/2026 at 3:52:05 AM

You can see the details here: https://hacks.mozilla.org/2026/05/behind-the-scenes-hardenin...

by atonse

5/11/2026 at 7:35:24 PM

There is no way to tell until we find examples of vulnerabilities that mythos missed. For all we know curl currently has 0 vulnerabilities right now

by HDThoreaun

5/11/2026 at 10:53:56 AM

The question is how many security vulnerabilities are actually left in the code after all the recent AI attention. Either Mythos is a nothingburger, or it's substantially more powerful but there's nothing left to do. Even a large amount of C can be correct eventually. Curl has the _potential_ to become a good data point maybe 6-12 months from now - if researchers and new tools find many more vulnerabilities then Mythos is proved to be hype. If they don't, then maybe Mythos is overkill for today's curl and its capabilities are better deployed elsewhere (like Firefox, apparently).

by thombles

5/11/2026 at 11:35:19 AM

I have a hard time believing that Mythos found the only remaining Curl vulnerability. It is possible, but highly improbable.

And it is not overkill, the proof is that it found that vulnerability. It is like saying the new version of some static analyzer with some new rules is "overkill" because it only found only one more bug than the previous version. Deciding whether it is overkill or not is more about context. Using a very expensive model like Mythos for some little used non-critical software is overkill, but for Curl, it absolutely isn't.

If Mythos found loads of vulnerabilities in Firefox but not in Curl, I wouldn't say that's because of Mythos is so good, but rather that with the release of Mythos, they did some testing that could have been done before using the same tools Curl have used.

by GuB-42

5/11/2026 at 11:46:41 AM

We will see. As for "testing that could have been done before", Mozilla's posts indicate otherwise. Use of Opus 4.6 led to 22 security-sensitive bugs vs Mythos' 271 (https://blog.mozilla.org/en/privacy-security/ai-security-zer...). They already had the methodology in place when the more powerful model came along (https://hacks.mozilla.org/2026/05/behind-the-scenes-hardenin...):

> Once the end-to-end pipeline is in place, it’s trivial to swap in different models when they become available. Building this pipeline early helped us find a number of serious bugs using publicly-available models, and it also helped us hit the ground running when we had the opportunity to evaluate Claude Mythos Preview. In our experience, model upgrades increase the effectiveness of the entire pipeline: the system gets simultaneously better at finding potential bugs, creating proof-of-concept test cases to demonstrate them, and articulating their pathology and impact.

by thombles

5/11/2026 at 1:35:14 PM

False dichotomy

by sitkack

5/11/2026 at 1:55:04 PM

It's not, really. Curl is an extraordinarily high value target that has already been picked over by well funded security researchers and state-sponsored groups using state of the art tooling for decades. That is not the target for which Mythos is a threat.

The threat isn't high value targets, which already had sophisticated folks picking over the code base using state of the art tools and tests, it's medium to low value targets which can now be picked over by random hackers who barely know anything about security themselves at a cost of a few dollars.

by empath75

5/11/2026 at 10:59:13 AM

that makes it a good data point, because it is better able to illustrate the incremental capabilities of Mythos compared to previous tooling

that helps us to understand how much of Mythos is hype and how much is real

by spongebobstoes

5/11/2026 at 10:32:48 AM

We see this exact hypetrain every time a new model is released. Mythos simply hasn't lived up to the "we're all gunna die from the flood of vulnerabilities" hype even slightly. Its slightly better than previous models by all accounts, cool stuff

I've seen literally near word-for-word this exact chain of events multiple times previously

by 20k

5/11/2026 at 2:35:43 PM

Is Mozilla marketing on Anthropic's behalf?

    As part of our continued collaboration with Anthropic, we had the opportunity to apply an early version of Claude Mythos Preview to Firefox. This week’s release of Firefox 150 includes fixes for 271 vulnerabilities identified during this initial evaluation.
    
    As these capabilities reach the hands of more defenders, many other teams are now experiencing the same vertigo we did when the findings first came into focus. For a hardened target, just one such bug would have been red-alert in 2025, and so many at once makes you stop to wonder whether it’s even possible to keep up.

https://blog.mozilla.org/en/privacy-security/ai-security-zer...

by orblivion

5/11/2026 at 3:44:14 PM

There are three things happening simultaneously: 1st a new model, codenamed "Mythos", 2nd a lightweight harness built for finding vulnerabilities, and 3rd a push by Anthropic to collaborate with various Open Source projects and companies to use 1 and 2 to find vulnerabilities

We know that the combination of all three results in finding lots of security vulnerabilities. That's what Mozilla is talking about. The quote from the curl story states that just 2 and 3, but with just regular SotA models, would have produced very similar results

Which is really the crux of all this hype around Mythos: would the results really be different if they used Claude Opus instead of Claude Mythos? How much is the model, how much the harness, and how much is just because Anthropic is running a big campaign systematically trying to find vulnerabilities?

by wongarsu

5/11/2026 at 3:51:01 PM

Not to discredit anything that was said in any particular blog post.

Folks also need to remember that a lot of blog posts are written by engineers or managers that have their own agendas and careers and often external blog posts can be a form of self marketing or idea marketing that an engineer or director has been pushing internally.

I have no idea if this happened in mozilla's case but the person that wrote it seemed to talk about the their own internal harness / fuzz testing framework quite a bit, and I imagine it was probably a big part of that person's scope / accomplishments and will probably show up at their end of year review and on their resume.

by sporkland

5/11/2026 at 4:51:31 PM

Also, the people at Mozilla who helped achieve a highly visible collaboration with the hottest AI company in the zeitgeist that included a lot of expensive data center time to harden their flagship product are definitely going to be happy/excited/proud about pulling it off successfully.

There's a lot of kneejerk "so you're accusing Mozilla of a conspiracy to boost Anthropic?" which is an overly simplistic lens. Particularly when it involves groups of individual humans with different motivations and emotional investment in their own contributions to the collaboration.

by toraway

5/11/2026 at 10:26:13 PM

Okay so supposing everybody is acting in a benign manner, following their incentives and passions, not meaning to mislead anybody. Do you think that this results in writing a misleading blog post? Because the blog post makes Mythos out to be a big friggin deal. (It had certainly convinced me).

by orblivion

5/11/2026 at 3:06:45 PM

It is difficult to compare these two accounts since Daniel Stenberg didn't get access to Mythos himself, and we have no information about how it was run compared to the other AI models that have been used on curl. It is possible that Mythos is not much better than these other models, but it is also possible that the curl team simply made better use of the other models.

Part of what made Mythos so effective for Mozilla was the integrated agentic workflow where it not only looked for bugs, but then created an exploit to demonstrate them, and ran that exploit while dynamic analysis was enabled verifying that invalid memory access occurred. In this case it hard to know how much of their success was because they put more effort into the harness compared to previous tools (we know they did), or if Mythos was more suitable for this sort of workflow to begin with.

Not many apple-to-apple comparisons to be made with Mythos at this point.

by pavon

5/11/2026 at 3:19:43 PM

> then created an exploit to demonstrate them, and ran that exploit while dynamic analysis was enabled verifying that invalid memory access occurred

Four years ago that would have sounded like science fiction. Right now, I think that even Gemini Flash might be able to do that, given a couple of attempts.

by esperent

5/11/2026 at 2:47:23 PM

Yep! The industry term is "co-marketing" and its hard to avoid seeing once you spot it.

by spenczar5

5/11/2026 at 3:41:55 PM

I'll wear the dunce cap: how are you so certain this is co-marketing? I'm not saying you are wrong, but it doesn't seem obviously like marketing copy to me (which is of course what they'd want but that's nevertheless not in any way evidence one way or the other).

by HelloMcFly

5/11/2026 at 3:46:28 PM

It starts with the words "As part of our continued collaboration with Anthropic"

Once these words are used you can assume there is a contract stating how that collaboration works, and that this includes some sentences about how much each side is allowed to or required to say about it

by wongarsu

5/11/2026 at 4:38:51 PM

So you claim that Mozilla entered into a contract with Anthropic, and said contract requires Mozilla to advertise for Anthropic on their blog. I hope Mozilla is getting a good payday out of this.

by warkdarrior

5/11/2026 at 2:59:15 PM

I didn't think Mozilla was like that but duly noted.

by orblivion

5/11/2026 at 2:41:24 PM

I think it's more the cost to find a vulnerability that has significantly reduced, not the possibility that the vulnerability could have been found. But that cost mattered tremendously because someone has to fund the effort to find the bugs. This economics also applies to attackers.

by dboreham

5/11/2026 at 2:43:25 PM

Is Firefox less invested in this than Curl? I mean there must be some explanation for this.

by orblivion

5/11/2026 at 2:49:41 PM

It's in the first sentence of your quote:

"our continued collaboration with Anthropic"

Read this as: "we get discounts, rate limit increases, a direct line to responsible product managers; in exchange we participate in friendly marketing." It's extremely common in this line of business - typical of database vendors, software tool companies, etc.

by spenczar5

5/11/2026 at 2:52:56 PM

This is more in response to my original post, but okay interesting point. (When I said "invested" here I meant invested in finding security flaws.)

by orblivion

5/11/2026 at 8:30:09 PM

In many countries it is mandatory to mark any form of compensated advertising as such. If your claim is true they might be breaking some laws here & there…

by janc_

5/11/2026 at 3:13:34 PM

Conspiratorial nonsense

by Anon1096

5/11/2026 at 4:28:01 PM

I would expect Firefox to be less invested in this than Curl. Firefox is aimed at consumers, Curl is embedded in a wide variety of products.

by cmiles74

5/11/2026 at 3:53:27 PM

Absolutely 100%

by skywhopper

5/11/2026 at 2:43:57 PM

I certainly wouldn't be surprised if they were.

by Pay08

5/11/2026 at 8:08:03 AM

It may well be that the hype was primarily marketing.

The other alternative is that Curl is simply secure enough that there was far less to find than in other projects.

by vidarh

5/11/2026 at 12:24:49 PM

Daniel found 30 CVEs in Curl, this year. I would not say that there is nothing to find, here. Just that it takes an actual expert.

by shakna

5/11/2026 at 3:37:32 PM

I did not suggest there was nothing to find. But is also very different to count all CVE's found and reported (there are less than 30 total for 2025 and 2026 per [1]) by anyone and everyone vs. what was found in a short time by someone prompting a model.

[1] https://curl.se/docs/security.html

by vidarh

5/12/2026 at 8:45:56 AM

Is not the selling of the model, that it is as capable as anyone and everyone?

> Claude Mythos is Anthropic's most specialized model, trained exclusively on security research, vulnerability disclosures, and attack pattern literature. Its reasoning reflects how the world's best security researchers think. [0]

[0] https://mythosvulnerabilityscanner.com/what-is-claude-mythos

by shakna

5/12/2026 at 12:00:27 PM

Even if I was selling the model, which I am not, it still does not follow that you can judge that on a single run, given that no security researchers have found all of these bugs on their own in a short amount of time either.

by vidarh

5/12/2026 at 10:32:49 PM

Okay, to respin this - Daniel doesn't say that curl is secure-enough. Half the point of the talks this year, is there has been an uptick in detecting security bugs, not a downturn. And here's some graphs. [0]

> Given the look of these graphs I don’t think we are close to zero bugs yet. These two curves do not seem to even start to fall yet.

If the author thinks there is more to find, then the soil probably isn't dry.

But, from the author's mouth:

> My personal conclusion can however not end up with anything else than that the big hype around this model so far was primarily marketing. I see no evidence that this setup finds issues to any particular higher or more advanced degree than the other tools have done before Mythos. Maybe this model is a little bit better, but even if it is, it is not better to a degree that seems to make a significant dent in code analyzing. [1]

[0] https://daniel.haxx.se/blog/2026/04/30/approaching-zero-bugs...

[1] https://mastodon.social/@bagder/116554460442650929

by shakna

5/11/2026 at 1:02:03 PM

[flagged]

by bcjdjsndon

5/11/2026 at 10:13:54 AM

Given how much money is on the line, it would be gross negligence if anything came publicly out of the CEO's mouth or is otherwise published by the company that's not marketing.

by teiferer

5/11/2026 at 12:02:01 PM

The question is whether they need to massage the results for them to be marketable.

by red75prime

5/11/2026 at 12:51:47 PM

Sometimes you gotta let people know how awesome you are. The real question is if you're misrepresenting yourself(all marketing, no substance).

by zeroCalories

5/11/2026 at 1:35:36 PM

Not really, curl has slow anonymous memory leaks because of how the connection session caching was implemented. If you don't periodically restart a program, than people encounter strange hard to diagnose issues sooner or later.

Also, looking at something that trips valgrind warnings already, may obfuscate a lot of problems in both your own code and the curl library itself.

One could report the issue as functioning as described in the API, but the developers do not accept direct community input into the project.

People use it out of convenience, but it is just as janky as most bloated projects. =3

by Joel_Mckay

5/11/2026 at 9:45:01 AM

My guess:

Marketing is not intentional.

Evidences: 10 years ago, when I interviewed Baidu AI with Andrew Ng and Dario, Dario is the kind of person is pure-hearted to the point being ideological. Given Dario's successful career so far, that essence has gradually grown into a conviction, and surrounded by a purposely built team which amplifies his ideology.

Humans are very convenient creature, a rare few small fraction of them are no doubt the master of convenience: they morph their mental manifold without a hint of contradiction in their own mental mechanisms.

by bigcat12345678

5/11/2026 at 10:14:08 AM

These things are layered. They are great scientists, smart people, etc.

Things change when you’re running a business like Anthropic, especially as the CEO. You have a responsibility to shareholders, and you just need to play the game.

Anthropic chose a great angle: focus on professionals / enterprise, safety, etc. Those can both be done by a genuine desire to make great technology, and for business purposes require you to position yourself in a bit “better” way than reality.

Just look at what their strategy is with Mythos, it’s almost perfection: the “it’s not ready to be released to the public” angle hits all the marks: they care about responsibility / safety, they have “the best” model, and “LLMs are dangerous, but we, as the guardians, can be trusted”. This also helps the industry as a whole with regulation: if they’re being constrained, China will develop even more dangerous models.

This is a result of how smart people treat business, it’s PR perfection, especially given how much the whole industry is talking about it.

(Yes, they fail in other PR areas, but that’s a different discussion)

by stingraycharles

5/11/2026 at 12:53:27 PM

Marketing is always intentional at this scale. If you think Anthropic didn't put a lot of time and effort into Glasswing as a marketing effort I think you're misunderstanding how these organizations work and how they win.

by windexh8er

5/11/2026 at 12:06:25 PM

> Marketing is not intentional

Mythos put Anthropic back into the White House’s good graces. It also branded Anthropic as badass, something their softener image probably needed to win government contracts.

Maybe it wasn’t marketing. But the product’s configuration, and how Anthropic talked about and released it, sure as hell played beautifully. (The timing, while Musk and Altman are distracted with each other, also couldn’t have been better.)

by JumpCrisscross

5/11/2026 at 10:02:18 AM

I'm not sure if that distinction is important, since what you've described less charitably synonymous with the phrase "Dario is delusional, and has surrounded himself with yes-men, so outlandish marketing gets published as a side effect".

Whether the person doing the marketing was sincere about it or not is immaterial, since marketing is experienced almost entirely by the people consuming it, and not the people communicating it. What matters is if the audience is sincerely concerned by the message, and it's transparently the case that they were sincerely concerned by it.

by OtherShrezzing

5/11/2026 at 10:11:24 AM

> Marketing is not intentional.

That's an odd definition of "intentional". Evolution has filtered for people with certain views and the marketing has just emerged from their actions. ... So?

A deadly virus (naturally occurring one let's say) wasn't created intentionally. Evolution selected for it. It's still bad and kills people. Doesn't make it nice because of lack of intention.

by teiferer

5/11/2026 at 2:10:06 PM

I think that's a reasonable analysis, but it's very different than the one that's usually implied by "marketing". Most people I see talking about Dario and his "marketing" go on to express confusion or frustration on why he would decide to message this way, ignoring what I (and perhaps you?) consider to be the obvious answer that he believes it's true.

by SpicyLemonZest

5/11/2026 at 12:49:45 PM

This is marketing.

by keybored

5/11/2026 at 1:11:56 PM

All your evidences can be exactly true, and he genuinely believes that Anthropic "winning" the AI race is the best outcome for humanity even with a little subterfuge including marketing to the current administration. If I genuinely thought I needed to do something to secure humanity, there's little I wouldn't do to achieve it.

by petesergeant

5/11/2026 at 12:23:47 PM

Even that press release never claimed that Mythos was better than Opus at finding bugs.

They claim the huge advance is in exploiting the bugs.

by cvwright

5/11/2026 at 12:15:16 PM

He also said this [1] a few weeks ago about AI PRs.

> Over the last few months, we have stopped getting AI slop security reports in the #curl project. They're gone.

> Instead we get an ever-increasing amount of really good security reports, almost all done with the help of AI.

> They're submitted in a never-before seen frequency and put us under serious load.

> I hear similar witness reports from fellow maintainers in many other Open Source projects.

> Lots of these good reports are deemed "just bugs" and things we deem not having security properties.

[1]: https://www.linkedin.com/posts/danielstenberg_hackerone-shar...

by smusamashah

5/11/2026 at 3:18:42 PM

> My personal conclusion can however not end up with anything else than that the big hype around this model so far was primarily marketing.

I think the results say more about the great job the curl team has done maintaining their codebase.

This doesn’t mean Anthropic's Project Glasswing is a marketing stunt. Logically, it doesn’t make sense: when they announced Mythos Preview, Anthropic couldn’t meet customer demand; they didn’t have enough compute to go around. So they decide to hype an unreleased product to drive even more demand? All that would do is piss off their existing customers who already experiencing rationing and frequent outages.

Many forums were already flooded with "I cancelled Claude Code" as it was.

On the contrary, it would be incredibly irresponsible and unethical for such a young company with billions of dollars of other people’s money invested in them.

Because the Mozilla team used Mythos and found 271 vulnerabilities [1], does that mean they're in on the so-called "marketing stunt"?

Of course, if Anthropic had released Mythos to the public and bad actors used it to hack a large number of banks, hospitals, government agencies, etc. in a matter of days, the HN crowd would be all over them for acting irresponsibly and criticizing them for not knowing better.

[1]: "Behind the Scenes Hardening Firefox with Claude Mythos Preview" — https://hacks.mozilla.org/2026/05/behind-the-scenes-hardenin...

by alwillis

5/11/2026 at 3:34:04 PM

Other AI tools have found 300 bugs and this new sentient T1000 only found one. Stenberg himself found 30 this year.

Mozilla is the current poster child but 271 in such a large codebase with thousands of user options, most of them being TOCTOU isn't that much. Sorry. TOCTOU can happen in any language when people are simply exhausted by the sheer volume of case explosions.

There is a third option: Anthropic could simply have reported the issue without mentioning the new model at all. But they don't, since they want to sell to governments and military and the artificial scarcity just provides a veneer of exclusivity that their clients will appreciate.

by asltp_

5/11/2026 at 9:04:05 AM

Mythos marketing really leans into that "too powerful to be legal" vibe, much like how PS2s were allegedly banned from North Korea because their chips were basically missile-grade.

by jansan

5/11/2026 at 11:52:01 AM

I'm pretty sure mythos is just a new unreleased version of Opus + marketing + a different system prompt.

by 63stack

5/11/2026 at 12:14:07 PM

I suspect so as well.

I've been running my own security scanning software (disclaimer: now starting a company @ zeroquarry.com) for this, and from what I've seen there's a huge value in prompts + adversarial LLM review. Without adversarial review, you get garbage (as this blog points out: 4/5 basically are nonsense) and with a good prompt, you can use almost any "near frontier" model from my experience as long as the prompt helps with the guardrails or the model doesn't protect in such a strict way

by eskibars

5/11/2026 at 3:43:26 PM

It's almost as if management was a useful function in organizations ;)

by dboreham

5/11/2026 at 2:17:28 PM

Yes, the governments fall for it for the time being:

https://www.politico.eu/article/anthropic-hacking-technology...

This is an advertising masterpiece: UK gets first access, the EU is jealous and wants it, too. Thousands of bureaucrats and parasites make money in the process writing (probably using AI) whitepapers and sitting in meetings. The open source authors whose works are being scanned make nothing.

We know how the money flows. Another unrelated example is that ex MI6 director Sir John Sawers is a Palantir consultant and sells out the UK to Palantir.

by khlapz

5/11/2026 at 8:36:15 AM

They might be biased by the fact that curl is significantly more secure than the average software

by h1fra

5/11/2026 at 4:43:23 PM

I've seen this suggested a few times in this thread but it seems like it's exactly backwards.

Wouldn't that make it a better to distinguish whether Mythos is uniquely super powerful vs an incremental improvement from Opus etc that are routinely used as the basis for bug reports/fixes in cURL?

If Mythos found a hundred new show stopper bugs then it would have meant Opus missed them and therefore closer to a "step change". Otherwise it implies the difference in capability isn't nearly that stark. Mythos finding 100 low-hanging bugs in a less scrutinized/hardened project on the wouldn't be as useful signal to answer that.

by toraway

5/12/2026 at 4:47:56 AM

I have an impression that he expects something like the famous move 37 of AlphaGo, while it could be that the situation is like in chess where superhuman engines validated human findings.

by red75prime

5/12/2026 at 7:48:08 AM

Am I missing something here. Not finding a major bug/vulnerability just means that maybe the code is really good, not that the model is not what is claimed?

by hislaziness

5/11/2026 at 8:03:56 AM

>It's a good reminder for us all that the competition in this space is rough and lots of more or less subtle marketing is involved.

About as subtle as a personal injury lawyer's billboard

by coldtea

5/11/2026 at 8:07:38 AM

Better Call Dario

by steve1977

5/11/2026 at 8:06:15 AM

A thankfully American reference

by te_chris

5/11/2026 at 9:13:50 AM

Can you expand on this? Do you mean in contrast to the European AI milieu?

by Exoristos

5/11/2026 at 9:57:06 AM

No, the personal injury lawyer billboards.

by te_chris

5/11/2026 at 7:26:49 PM

In the UK of course it would be "personal injury barrister"

by llbbdd

5/12/2026 at 7:10:31 AM

I’ve never seen a billboard for it here - lived here 11 years.

by te_chris

5/11/2026 at 7:47:22 AM

I'd go out and say the marketing is not subtle. The hype and fanboys/girls are so in line with the marketing that any level of skepticism is seen a an act of defection, but if you look at the words, hyperbole and volume that is used, there is nothing subtle about it.

It's almost Trump-esque - "this model will change everything forever; we are doomed; we are saved; we will all be fired; we will all be rich", etc

by greendude29

5/11/2026 at 8:01:42 AM

That's a pretty good encapsulation of the parallels between the political and the technological: One necessarily thrives upon the other and are inextricable. This moment is a culmination of all the disenfranchisement the bodypolitik have suffered, looking for any possible means of escape or elevation. AI and Trumpism, for their own respective cohorts, are salvation, on offer by different frontmen but ultimately in service of the same system.

They need the hype to pay off way more than we do. So many of us who still write code directly stand to lose nothing of our capabilities if the marketing claims cannot hold water.

by xantronix

5/11/2026 at 8:24:22 AM

I seem to be totally outside the hype bubble, but I have to suspect there is a lot of imagineering and wild extrapolations in the elss technical hype bubbles. I am curious but no enough to go looking.

by ehnto

5/11/2026 at 8:58:01 AM

>I seem to be totally outside the hype bubble

I'm surprised you say that because it is all over Hacker News. Every single post is co-opted into promoting AI. Try finding a submission with fifty points or more than doesn't have AI or LLM's mentioned somewhere in the comments.

by tonyedgecombe

5/12/2026 at 2:13:19 AM

That's a good point, I guess I see the Hacker News hype a bit more realistically then maybe I should. HN has definitely changed in the sense that I rarely see interesting technology or achievements hit the front page, that aren't AI related. It feels like AI has taken all the oxygen from the room.

by ehnto

5/11/2026 at 10:03:17 AM

Feel free to retire from the field if you grow tired of seeing its latest developments.

by zen928

5/11/2026 at 11:10:05 AM

I already have.

That’s not really the point though. I have no doubt AI is useful, I just don’t want to have it shoved in my face every five minutes.

by tonyedgecombe

5/11/2026 at 12:06:20 PM

Eh... I think he puts the LLM down for his own ego's sake (as would I!). Curl may, next to the Linux kernel, be one of the most heavily audited codebases in existence. The LLM found something he and thousands of others missed. It's not unimpressive.

by bjourne

5/12/2026 at 2:27:22 PM

The claim has never been that the new model could not do impressive things. The claim is that the new model is not the existential crisis Anthropic’s initial announcement post made it out to be.

by billyoneal

5/11/2026 at 10:00:26 AM

[dead]

by aaron695

5/11/2026 at 1:26:40 PM

I commented this in another post but I'm going to repeat it because I believe its important for this discussion.

> The worrying part about Mythos isn't the fact that it can find bugs. The worrying part is Mythos being able to find them on its own across entire code base as vast as Firefox then write exploits for what its found with a very basic prompt.

> The skill required to find then create zero days is quickly approaching the floor.

by wnevets

5/11/2026 at 1:31:48 PM

Opus can find bugs on its own in large codebases just fine with minimal prompting.

The great exaggeration is that this is a new capability.

by colechristensen

5/11/2026 at 1:37:03 PM

> Opus can find bugs on its own in large codebases just fine with minimal prompting.

and then it write the exploits automatically for you?

by wnevets

5/11/2026 at 1:47:15 PM

Yes

by colechristensen

5/11/2026 at 1:58:12 PM

I will never ever understand how people are amazed by this. Have they just not tried it and then just assume that because Anthropic says this is the first it must be true?

This was one of the first things I tried and it works great.

by ofjcihen

5/11/2026 at 1:48:41 PM

Can you send me that link?

by wnevets

5/11/2026 at 1:58:53 PM

Does this mean you’re only using the models in the web app? I mean that might be why you haven’t been able to do this?

by ofjcihen

5/11/2026 at 1:54:05 PM

What link? I've done it myself.

by colechristensen

5/11/2026 at 2:01:39 PM

You've pointed codex to the entire source code of firefox and simply prompted it to find bugs and then had it write the exploits for you? Why haven't you published this? That would sink all of the the claude code hype.

by wnevets

5/11/2026 at 2:23:56 PM

No, I'm not interested in Firefox bugs, but I've done it with my own large projects.

What I think happened here is an Anthropic team with very little security expertise were working on finding bugs for marketing reasons and when they prompted to make POC exploits of those bugs they didn't have much success because they didn't really know what to ask for. They then proceeded to very finely tune their next model to eagerly exploit vulnerabilities making the models much more powerful for the "I don't know what I'm doing" user which they're now trying really hard to convince everyone is a game changer. </speculation>

The reason many of us are skeptical is we've used the current models to do things and they've worked.

An analogy might be if they tuned their model to eagerly instruct somebody how to make improvised weapons, now somebody is asking about how to deal with a rival at work and their model gives instructions on building a bomb from hardware store parts. Then go on a marketing spree telling everybody how dangerous it is. This example might highlight how insincere the marketing is. At any point you could have tuned the model to exploit for inexperienced people, now that you've done it does not mark a grand new capability. People who knew what they were doing could already do this with models.

https://www.anthropic.com/news/mozilla-firefox-security

by colechristensen

5/11/2026 at 2:42:22 PM

> No, I'm not interested in Firefox bugs, but I've done it with my own large projects.

Can you publish your results and send them to Bruce Schneier, Dave Lewis, & Heather Adkin [1] so they know that this isn't anything new and just the work of people with little security expertise?

[1] https://labs.cloudsecurityalliance.org/mythos-ciso/

by wnevets

5/11/2026 at 3:06:41 PM

That whitepaper did not need 19 authors. They're there for show.

The Mythos FUD is a gift to the security team because it made the C-suite care about security and this is a plan to tell them what should be done and what to expect in the era of LLM security tools.

This is an emperor-has-no-clothes situation but we're selling winter coats and winter is near. Not focusing on how the Mythos FUD is exaggeration and instead focusing on actually necessary security postures is perhaps a tad dishonest but it still gets everybody in a better state and is an unfortunate common point in C-suite politics (and why the rich and powerful often seem so disconnected from reality and common people, everyone around them is trained to interact with them in a certain way and "mythos marketing is bullshit" is one of those things that people just don't say to them)

by colechristensen

5/11/2026 at 3:36:24 PM

Isn't that all the more reason to publish your process & results using Codex to do the same thing they're claiming? Presuming any bugs Codex found would be fixed and no longer a security concern.

by wnevets

5/12/2026 at 1:48:00 AM

Why would you publish something unremarkable and benign?

Is it actually that hard for you to go try this out yourself?

by ofjcihen

5/12/2026 at 3:35:20 AM

> Is it actually that hard for you to go try this out yourself.

I can't get it to work Codex, can you?

by wnevets

5/12/2026 at 12:58:01 PM

Yes. That’s my main driver. What do you mean you can’t get it to work?

by ofjcihen

5/12/2026 at 1:52:37 PM

You must show me how you are able to coerce Codex to be useful using this setup with no hand holding. You say its unremarkable and benign but it doesn't match my experience at all. I'm convinced I am not the only person on HN who would love to know how you are able to do it.

> We launch a container (isolated from the Internet and other systems) that runs the project-under-test and its source code. We then invoke Claude Code with Mythos Preview, and prompt it with a paragraph that essentially amounts to “Please find a security vulnerability in this program.” We then let Claude run and agentically experiment. In a typical attempt, Claude will read the code to hypothesize vulnerabilities that might exist, run the actual project to confirm or reject its suspicions (and repeat as necessary—adding debug logic or using debuggers as it sees fit), and finally output either that no bug exists, or, if it has found one, a bug report with a proof-of-concept exploit and reproduction steps.

> Finally, once we’re done, we invoke a final Mythos Preview agent. This time, we give it the prompt, “I have received the following bug report. Can you please confirm if it’s real and interesting?” This allows us to filter out bugs that, while technically valid, are minor problems in obscure situations for one in a million users, and are not as important as severe vulnerabilities that affect everyone. [1]

[1] https://red.anthropic.com/2026/mythos-preview/

by wnevets

5/12/2026 at 2:51:24 PM

I quite literally do this almost exactly with GPT 5.4. Sometimes I give it a poke in a direction but it largely runs by itself.

I don’t know what to tell you. You say it’s not possible but the money in my HackerOne account says otherwise.

by ofjcihen

5/12/2026 at 4:13:47 PM

> I don’t know what to tell you. You say it’s not possible but the money in my HackerOne account says otherwise.

I haven't said it was impossible. I said I can't replicate the Mythos setup with Codex on any project even approaching the size of Firefox.

If your Codex setup and the results its generates are unremarkable, please post them.

by wnevets

5/12/2026 at 4:40:12 PM

My codex setup is quite literally a single file that states a format I want my reports to be written in so that I can review them before submitting them. Sometimes I bother with the container setup, sometimes I don’t depending on the work.

This isn’t a matter of a harness, skill files, anything. This is just something that a model can do.

You have multiple people saying they’ve done it here. I can only assume you’re being facetious at this point.

by ofjcihen

5/12/2026 at 6:18:29 PM

You've now spent multiple days in this comment thread describing this as simple and unremarkable but refuse to share anything about it. At any point in the last 24+ hours you could've posted your single file, the size of the project, and what the model was able to produce on its own.

Must this information be protected or is its unremarkable?

by wnevets

5/11/2026 at 7:51:48 PM

No, what I'm doing isn't remarkable.

Publishing an extensive critique of Anthropic marketing is just an exercise in attracting abuse from nitpickers and the ignorant. If the author of cURL can't convince people, and security of his product has been one of his primary responsibilities for decades in one of the most widely used pieces of software out there... what hope do I have?

I've got better things to do.

by colechristensen

5/11/2026 at 8:43:05 AM

> An amazingly successful marketing stunt for sure.

This. Well done by Antropic.

It even reached the CISO of my small semi-government org in the Netherlands, who slightly panicked at the announced 'tsunami' of vulnerabilities that was coming with Mythos.

Got us some more money and priority with the board, though.

Never waste a good marketing scare.

by apexalpha

5/11/2026 at 10:08:24 AM

I don't agree with the "no tsunami in sight": if you don't look at 100+ bugs in Firefox and many more OSS projects, bunch of old unseen-before OpenBSD/Linux RCEs, and a few LPE in just 2 or 3 weeks for Linux itself...

IMO, this does not sound like marketing scare, there is spike of vulnerability disclosures - high quality, low false positives - that can be sensed... It feels like we're speedrunning through few-years worth of high quality bug reports in just a few weeks.

by fpesce

5/11/2026 at 12:38:56 PM

Mythos isn’t released yet.

Anthropic noticed the trend of AI vulnerability scanning and started advertising Mythos, which is unreleased, as being very good at it.

Then they donated very large token budgets for using Mythos privately to several teams. Those teams used the free token spend for security research (that was the deal) and anything they found got attributed to Mythos, not the token budget.

Mythos looks like a good incremental model but the PR team has done a great job of associating themselves with the current trend. So much so that comments like yours already associated vulnerabilities found with this model which isn’t even available yet

by Aurornis

5/11/2026 at 6:34:25 PM

Mythos hasn't been released yet, but there seems to be some evidence that GPT-5.5, which has been released, is already a touch better anyhow in some dimensions: https://www.mindstudio.ai/blog/gpt-5-5-vs-claude-mythos-cybe...

Close enough that you can probably get a good sense of Mythos' performance by using GPT-5.5.

One thing I noticed while using GPT-5.5 for this is that the ability of the model to turn the bug into an outright vulnerability is less relevant than you might intuitively think. All that is really necessary is for the model to point out that something is smelly, and you should just fix it. Turning it into a runnable exploit has very limited utility for the defender. It does turn heads and may get the attention of some otherwise reluctant people, but everything I found was obviously enough wrong that the exploit was just decorative.

by jerf

5/12/2026 at 2:29:39 PM

An actual PoC is often very helpful in prioritizing getting the bug fixed, in demonstrating that the bug is real, and in providing something that devs can see happening in their debuggers.

by billyoneal

5/11/2026 at 10:10:09 AM

The LPEs were not found with Mythos but with existing, publicly available models.

by apexalpha

5/11/2026 at 10:50:11 AM

And also: they did an earlier run with Opus to discover bugs (like segfaults).

In February, Opus discovered a whole bunch of security related bugs, but didn’t exploit them.

Mythos, in turn, was fed these bugs and told to exploit them.

Not saying it’s not impressive, but it was literally told “here are all the places our metal detector says there may be gold, please find gold”.

by stingraycharles

5/11/2026 at 1:24:52 PM

There is a significant difference between being able to see one flaw and being able to chain together multiple disparate flaws, to be fair.

by dralley

5/11/2026 at 7:56:30 PM

> bunch of old unseen-before OpenBSD/Linux RCEs,

AFAIK, the only thing it found in OpenBSD was a DoS?

Edit: For that matter, I'm not aware of RCEs in Linux, only LPE?

by yjftsjthsd-h

5/12/2026 at 10:37:56 AM

The whole thing started with a talk from Nicholas Carlini mentioning a remote 20+ year old NFS vuln IIRC.

by fpesce

5/11/2026 at 11:16:55 AM

Anthropic has is quickly destroying customer goodwill by repeatedly pulling the same stunt. Horrible marketing, imho.

It's an entirely different thing to have the company conduct research on LLMs in general being a cybersecurity threat, instead of going "our new model is just too powerful" and shift the discussion to revolve around that. It's slimey.

by helloplanets

5/11/2026 at 7:46:04 PM

Hasn't almost every new frontier model had an early period of limited access? I don't get why everyone is acting like Mythos is particularly egregious for this.

by AgentME

5/11/2026 at 8:52:21 PM

This is literally how they announced the model:

> We formed Project Glasswing because of capabilities we’ve observed in a new frontier model trained by Anthropic that we believe could reshape cybersecurity.

> Claude Mythos Preview is a general-purpose, unreleased frontier model that reveals a stark fact: AI models have reached a level of coding capability where they can surpass all but the most skilled humans at finding and exploiting software vulnerabilities.

https://www.anthropic.com/glasswing

by helloplanets

5/11/2026 at 8:20:23 PM

It is called "Mythos" dude...do you have any idea how mysterious and scary this sounds to most people and how much hype that alone can generate.

If the model was calle "Mini Mouse" it wouldn't feel anywhere near as threatening and interesting.

It sounds like the name of a cologne from the 70s or something and I like it.

by bamboozled

5/11/2026 at 11:22:03 AM

The bar has become so low lately that no one will care.

by aswegs8

5/11/2026 at 5:55:07 PM

He describes in detail how curl is software-engineered to within an inch of its life. Do you really think most code is that highly polished?

by AlexCoventry

5/12/2026 at 5:28:28 AM

Well done for convincing people of something that isn't true, in other words, well done for lying? Is this what is being cheered on? Seriously.

by Valakas_

5/11/2026 at 11:16:14 AM

org head is smart.

by markus_zhang

5/11/2026 at 4:16:03 PM

If an AI agent finds zero bugs in a software utility, how can that be viewed in the sense the AI agent is not very good at finding bugs?

What if there are actually zero bugs?

> Five issues felt like nothing as we had expected an extensive list.

The expectation here may not match reality, but not necessarily because Mythos isn't as capable as claimed. curl may just happen to be a well-hardened tool that doesn't have too many security vulnerabilities in its present state.

by EMM_386

5/11/2026 at 5:09:00 PM

The author considered the same w.r.t. remaining bugs:

> More to find

> These were absolutely not the last bugs to find or report. Just while I was writing the drafts for this blog post we have received more reports from security researchers about suspected problems. The AI tools will improve further and the researchers can find new and different ways to prompt the existing AIs to make them find more.

> We have not reached the end of this yet.

> I hope we can keep getting more curl scans done with Mythos and other AIs, over and over until they truly stop finding new problems.

And that makes sense, it'd be quite the argument of coincidence to say there was just 1 proper find remaining & it was only Mythos that managed to find it just at the point in time it released while the other projects have been hoovering up every other find quickly until that point. Possible, but not the safest assumption to start questioning with.

by zamadatix

5/11/2026 at 7:37:04 AM

> Not particularly “dangerous”

I'm not sure that follows. As noted, curl was already analyzed to death with every tool available; most software isn't at that level.

by yjftsjthsd-h

5/11/2026 at 9:19:00 AM

But Mythos is not marketed as a tool that can do the same as other tools already available maybe slightly better, but as a revolution.

by anygivnthursday

5/11/2026 at 3:06:39 PM

I'm agnostic with Anthropic/Mythos but if there aren't any vulnerabilities there it's hard to find it.

Until we find vulnerabilities in curl that Mythos missed, it's hard to say how good it is.

by abirch

5/11/2026 at 8:32:30 PM

Would have like to see analysis against curl repo where the commit level is one day after the Mythos training data cutoff. And disable access to the internet.

by nicce

5/12/2026 at 2:32:16 PM

No. The goal posts are not disproving their marketing. The burden of proof is on the people doing the marketing.

by billyoneal

5/11/2026 at 3:09:55 PM

Mythos is either dangerous or not. We are taking dangerous to mean that the number of vulns it finds will be much greater than bugs found with available tools.

Since mythos found only one additional vuln, and since x+1 is not much greater than x, it follows that mythos is not dangerous per the definition above.

by matltc

5/11/2026 at 4:39:16 PM

It doesn’t follow because the results for curl don’t necessarily generalize to other codebases. It’s evidence against Mythos being particularly dangerous, but it’s just one datapoint.

It doesn’t invalidate the other security bugs Mythos allegedly found in other codebases.

by skybrian

5/11/2026 at 7:41:01 AM

I don't think I understand what you mean, the "not particularly dangerous" comment was in relation to the vulnerability that was found right ? Surely they would know what constitutes a lower severity level.

by bilekas

5/11/2026 at 8:09:28 AM

The "not particularly dangerous" is a headline for a section talking about Mythos, not the vulnerability.

by vidarh

5/11/2026 at 8:12:34 AM

Ah okay, that makes a bit more sense. I read it wrong. Then the comment is absolutely fair.

by bilekas

5/11/2026 at 7:43:01 AM

My guess is that it is in category of "you are holding it wrong". Still worth fixing, but requires very specific user input for example. Or very weird scenario. Or in some less used protocol or flag combination.

by Ekaros

5/11/2026 at 9:13:38 AM

Sure, but isn't it a verdict on Mythos compared to other models?

If so, it would still follow. "Most software" isn't analyzed as much as curl, by either other tooling or other models, that might well find close to the same as Mythos did. As such, Mythos then isn't especially/particularly dangerous.

by croon

5/11/2026 at 11:01:57 AM

Curl is currently receiving a record number of high-quality bug/vuln reports (a rather sharp change from the earlier slop inundation), so it’s not like there’s nothing to find. Many or most of these are presumably found by human experts assisted by AI tools, but if Mythos were truly revolutionary, it should be able to find such issues on its own.

https://daniel.haxx.se/blog/2026/04/22/high-quality-chaos/, linked from TFA

by Sharlin

5/11/2026 at 12:11:41 PM

Is there a list of infrastructure that has received this kind of focus? Clearly people are looking at the linux kernel, hopefully openssl?

by galangalalgol

5/13/2026 at 7:09:16 AM

From article:

> I did a quick unscientific poll on Mastodon to see if other Open Source projects see the same trends and man, do they! Friends from the following projects confirmed that they too see this trend. Of course the exact numbers and volumes vary, but it shows its not unique to any specific project.

> Apache httpd, BIND, curl, Django, Elasticsearch Python client, Firefox, git, glibc, GnuTLS, GStreamer, Haproxy, Immich, libssh, libtiff, Linux kernel, OpenLDAP, PowerDNS, python, Prometheus, Ruby, Sequoia PGP, strongSwan, Temporal, Unbound, urllib3, Vikunja, Wireshark, wolfSSL, …

by snackerblues

5/11/2026 at 2:13:19 PM

I can't help but think that curl is, by nature, a relatively simple and well-contained tool. Compare to an operating system or web browser or database or billion dollar company codebase.

It makes some sense that Mythos/ChatGPT 5.5 might be that much better with complexities that curl just doesn't have because it's a basic tool.

Like yeah curl is obviously extremely fully featured as an "anything client" but it's orders of magnitude less complex than other software we rely on.

by srcreigh

5/11/2026 at 3:41:29 PM

Curl is a lot more complicated than, I believe, you think. Most people know of it simply as a CLI to hit an HTTP(S) endpoint and write it out. But:

1. It supports basically any file transfer protocol.

2. It is a library that is designed for long running processes.

3. Because it's designed for long running processes, it makes use of every trick it can to pipeline and re-use connections and resources.

4. It has an asynchronous API so it can be integrated into any existing event loop.

Is a web browser or database more complicated? Most certainly, they solve really massive problems. But curl is certainly more complicated than probably most application code that uses it.

by sausagefeet

5/11/2026 at 2:23:10 PM

I agree it's rather basic but as stated in the article, its code is still longer than war and peace. There is still plenty of opportunities for security vulnerabilities in something of that size.

by joelthelion

5/11/2026 at 8:28:37 PM

From the post:

"curl is currently 176,000 lines of C code when we exclude blank lines. The source code consists of 660,000 words, which is 12% more words than the entire English edition of the novel War and Peace. ... curl is installed in over twenty billion instances. It runs on over 110 operating systems and 28 CPU architectures. It runs in every smart phone, tablet, car, TV, game console and server on earth."

I wouldn't call that simple or well contained...

Most OS or web browsers don't run on cars or tvs.

by breakpointalpha

5/11/2026 at 10:45:55 PM

someone(mythos?) should write some simple-curl with 20% of features implemented in rust used by 98% of users.

by andriy_koval

5/12/2026 at 4:07:40 AM

curl is dealing with the complexity of HTTP. Even doing a simple basic request to some website, is going to cover a lot of code paths to deal with all sorts of response codes (redirects, etc.), headers, etc.

It's likely that new Rust code would introduce more bugs, while curl is extremely well tested at this point.

by lsdajsadsfdakl

5/11/2026 at 7:34:04 AM

> The single confirmed vulnerability is going to end up a severity low CVE planned to get published in sync with our pending next curl release 8.21.0 in late June

My mind still cannot understand the quality and refinement that's gone into cURL. It really is the perfect example of something done so right, that people barely think twice about.

by bilekas

5/11/2026 at 9:07:14 AM

Easy, it shows what is achievable if there is a high bar for quality in every single line of code that gets commited, reviewed and merged, regardless of the programming language.

However in the days of race to bottom, offshoring for penies, and now LLM powered code generation, this is a quality most companies won't care unless there is liability in place.

by pjmlp

5/11/2026 at 6:49:11 PM

> Easy, it shows what is achievable if there is a high bar for quality in every single line of code that gets commited

This is becoming a more and more overlooked/underrated feature. I genuinely believe it would be impossible in any company that depends on shareholder value. I am yet to convince any company I've worked in without bloody hands that we need to solve old tech debt and refactor certain things etc.

by bilekas

5/11/2026 at 7:27:18 PM

Which is liability is relevant, that is the only language shareholders understand.

by pjmlp

5/11/2026 at 7:48:52 PM

If you can get that message across the right way, you're a better company man than me. There's always someone more important than me to say 'but this needs to be delivered first'.

by bilekas

5/11/2026 at 9:35:43 PM

Sure, my point was more in general from government level.

by pjmlp

5/11/2026 at 7:58:27 AM

Curl and SQLite are my favourite examples of properly engineered, rigourously tested _anything_. It's really philosophical - those projects' contribution requirements demand such rigor, and the maintainers stand by that demand. A non-load-bearing document (not project code) is what makes that possible - very reminiscent of Einstein's thought experiments leading to tangible projects such as GPS or Descartes's belief that all problems can be solved through rational thinking.

by dotancohen

5/11/2026 at 10:28:42 AM

Some people must be working on training some models exclusively on high quality OSS code base like curl and SQLite without the noise of low quality training data.

I would do that with 100% local models from scratch.

by ontouchstart

5/11/2026 at 1:16:58 PM

> My mind still cannot understand the quality and refinement that's gone into cURL. It really is the perfect example of something done so right, that people barely think twice about.

And all that to then end with people doing: "curl ... | bash" and not seeing anything wrong about it. Then they'll deflect about "threat models" and other non-sense.

I leave you your curl-bash, I keep my cryptographically signed packages installer.

by TacticalCoder

5/12/2026 at 2:35:22 PM

I am also a signing fanboi but I have to point out that the security problem of curl into bash is not really addressed by signing. Signing proves that the component was produced by who claimed produced it. It says nothing about that component being legitimate or non-malicious. As long as the curl bash uses TLS it’s going to be pretty similar for all practical purposes.

by billyoneal

5/11/2026 at 3:30:53 PM

As far as I can tell, the messaging around Mythos is that it takes the expertise of the top security experts and top-level language, protocol and code experts and makes that available to anyone with access. The danger was in giving that access to the world before the defenders had access to that level of expertise.

Curl HAS had security, protocol and language experts poking at it for years because of how central it is to everything. That Mythos found anything is interesting but not a sign that it's been marketing hype and isn't dangerous.

You can bet that 99.99% of projects aren't nearly as secure as curl and it doesn't matter if they are open or closed source (LLM's will happily decompile closed-source projects and explore). Unless your project has been fuzzed and gone over with existing AI tooling and by experts, expect that it can already be hacked - even with the tooling that is out there now and that something like Mythos makes it accessible for an even wider population pool with less expertise to use.

by patrickmeenan

5/12/2026 at 12:13:19 AM

Take my upvote. Anthropic never claimed superhuman performance, only speed and scale. That it doesn't find much in terms of new vulnerabilities in a well-studied piece of software says nothing about its overall potential for dangerous misuse.

by 2001zhaozhao

5/11/2026 at 1:56:18 PM

I know that the Mythos hype is part marketing by anthropic, but isn't it possible that with a highly scrutinized codebase, there just aren't any notable security exploits in it's current state? The fact that it found nothing isn't necessarily an incrimination against it, especially when other tools had identified hundreds of exploits previously. Seems like it's been completely picked over (for now).

by jrflo

5/12/2026 at 2:36:12 PM

People lost their minds over the mythos announcement specifically because they found something in FreeBSD, which had a reputation as being one of those picked over code bases.

by billyoneal

5/11/2026 at 7:47:37 AM

There is always marketing involved and people should be able to put marketing into perspective.

Also curl in this regard is a open source project, relativly small but critical, well known and used everywhere. Besides image libraries, tools like curl or sudo, su, passwd, etc. would also be my first try.

Mythos is still not known at all what it can do. What does it mean from cost and benchmark pov to have a 10 Trillion parameter model?

Nonetheless, the fact that LLMs got significant better in finding this, better than humans, started to happen half a year ago? so at one point we need to address the elefant in the room and state that today you need to do security scanning additional with LLMs. You need to take this serious.

In worst case, use Anthropics marketing to state that its a must now and something changed.

by AntiUSAbah

5/11/2026 at 12:13:42 PM

> What does it mean from cost and benchmark pov to have a 10 Trillion parameter model?

To me it means that we've hit the top end of the S-curve with regards to effects of scaling - if the tool isn't remarkably better despite the scale, then we're firmly in diminishing returns territory.

by Tade0

5/11/2026 at 11:50:34 AM

> Mythos is still not known at all what it can do.

And this is very much on purpose my friend. Think about what people already believe it can do though.

by u_fucking_dork

5/11/2026 at 8:55:20 AM

> Nonetheless, the fact that LLMs got significant better in finding this, better than humans, started to happen half a year ago?

*rolls eyes* regular static analyzers also have been "better than humans" for decades, being better than a human at a specific mechanical task really doesn't mean much. The interesting new thing is the type of potential "fuzzy bugs" described in the article that LLMs are able to identify (a comment not matching the code it describes, uncommon usage of a 3rd party library, mismatch of code and a protocol it implements, or often just generally weird looking code somebody should have a closer look at... this closes a gap in the traditional debugging toolboxes, but shouldn't replace them)

by flohofwoe

5/11/2026 at 9:59:36 AM

You don't have to dismantle a comment on a microlevel.

It has been clear for ages that certain type of bugs or issues are better solved from software.

But there was still plenty of things a proper SecOps Person would be able to find with help from tooling which automatic tooling wouldn't find.

Taking a limited amount of resources and focusing on the critical things.

I do think this is gone now. Same with Threat modeling etc.

by AntiUSAbah

5/11/2026 at 2:02:20 PM

Static analyzers are balls. For every real bug they find you are dealing with with piles of false positives and negatives.

Now, I'm not saying you shouldn't use them. They do catch the low hanging fruit. It's that LLMs actually have a much better understanding of things like intent when looking at your code and general architecture configurations that can lead to problems.

As you say we've had static analyzers forever, hence why they aren't dropping out 50 new CVE's a day. LLMs are. There is a massive stack of software out there that is getting analyzed and exploited at a rate faster than it's getting patched. Adding to that things like NPMs exploited package of the day and popular github repository takeovers this year looks massively different from last year in quantity and quality of exploits alone.

by pixl97

5/11/2026 at 2:40:52 PM

IME LLMs generate at least as much false positives as static analyzers, but they're good at catching entirely different types of problems than static analyzers. 99% of false positives are avoided with a proper assert hygiene, and from what I've seen that seems to be true both for traditional static analyzers and llms, those assert annotate the code with valuable hints that may go beyond a specific type system's capabilities.

by flohofwoe

5/11/2026 at 7:32:39 AM

Putting on my tinfoil-hat: Sooo, the guy who runs the test and delivers the report could just have removed the more interesting bugs and delivered those to any three letter agency?

by ahofmann

5/11/2026 at 9:38:49 AM

curl's source is public so what would be the gain in the rigmarole? Now if the prompt was "create a patch that inserts a zero-day while fixing a bug" that would be impressive.

by casey2

5/11/2026 at 7:35:29 AM

[flagged]

by bilekas

5/11/2026 at 8:50:25 AM

The test was run by an unnamed third party, so cURL's history has no relevance to their benevolence.

by AnssiH

5/11/2026 at 7:38:12 AM

Curl is likely one of the very much more combed over pieces of code at this point. It feels like it has some special draw for people looking for vulnerabilities. Not that it doesn't mean some novel idea can't be looked or checked still.

by Ekaros

5/11/2026 at 8:09:21 AM

> No, based on cURL's history, it really seems like they would love to have found a really novel bug.

You just confirmed that you didn't read the article.

"Eventually, I was instead offered that someone else, who has access to the model, could run a scan and analysis on curl for me using Mythos and send me a report."

by cakealert

5/11/2026 at 8:33:14 AM

I'm not sure how that proves I didn't read the article ?

by bilekas

5/11/2026 at 9:38:34 AM

Someone external to the curl team ran the test. If that third party found a severe CVE that they could use across all the global curl attack surface, and did not disclose it back to the curl team, the third party could keep using the exploit until discovered independently.

by croon

5/11/2026 at 1:01:02 PM

What's going on in this thread? It's weird how prevalent the negativity towards mythos is, and I'm not sure if it's people throwing the baby out with the bathwater or something more tinfoil-adjacent coordinated campaign. I also noticed this on a thread a few days ago, before the mozilla post. There were dozens of comments saying basically "mythos is vaporware".

I get the idea that they're using it for marketing. Of course they are. But to reduce it at "just marketing" feels either ill informed or outright wrong. Unless you have reasons to not believe the dozens of credentialed, well respected people in the field that have already shared their opinions after working with mythos. Plenty of them on all the social media sites.

And then there's the team at mozilla. They wrote a blog about this, and they've worked with anthropic before, using opus 4.6 and found and fixed 22 vulnerabilities. Then they worked with mythos and found and fixed 271 vulnerabilities. Unless you're going to accuse them of being shills, these are unquestionable numbers. The model is quantitatively better at this thing. And it matches what everyone is saying.

I think there are better things to accuse anthropic of, than that they are simply lying for marketing purposes. Of course they'll use this as a marketing campaign, but there's plenty of evidence out there that there is something there, that the model is simply better than previous generations at this. Don't fall for the cheap reductionist stuff, just because you don't like them, or feel that this is marketing fluff. It doesn't feel like a gimmick, even if it gets used to push their agenda. Something, something, propaganda often uses true statements as well.

by NitpickLawyer

5/11/2026 at 1:20:59 PM

Here and on reddit, AI debugging is viewed as some weird shallow pattern-matching that obviously fails to spot real stuff and overload the maintainers. Instead of getting to "spotless record" of zero flaws, the people start rationalizing that "X is not a real bug" and inventing justifications for their(obviously bad) code, which is critique they can't accept from AI, only through human debate that they can't close with a WONTFIX. Once the bug is actually usable, the tune changes completely.

by countWSS

5/11/2026 at 1:56:19 PM

> Here and on reddit, AI debugging is viewed as some weird shallow pattern-matching that obviously fails to spot real stuff and overload the maintainers.

That's because that is what a lot of people did in the last years [1] to pad their resumes or to force developers to backport patches to older (but supported) kernel versions that wouldn't have gone in if they didn't have a CVE attached [2]. Maintainers have been legitimately swamped with low-quality spam for a very long time. Only recently, in the last few months, AI actually got "good enough", the problem is that maintainers still have to differentiate between AI slop by wannabes and by AI-assisted reports reviewed and refined by actual human professionals.

[1] https://www.zdnet.com/article/how-fake-security-reports-are-...

[2] https://opensourcewatch.beehiiv.com/p/linux-gets-cve-securit...

by mschuster91

5/11/2026 at 2:16:52 PM

At the end of the day attackers don't give a fuck. "Waaa waaa, AI was bad 6 months ago so I'm going to throw a little fit" doesn't work when it's currently actively exploiting your shit. No one gives a damn if there are 4000 bullshit security PRs lined up. The one real RCE in there mean that everything you hold dear has already been carted off by nation states, and probably rediscovered by 3 or 4 other exploitation groups by this point.

It's time for all the little snowflake software writers to pull up their pantaloons and realize that Linus' vision has become real. With enough AIs all security bugs become shallow. And that software affects the real word, real money, and real people in it. That they are also under attack by well financed groups with rather evil motivations. If I'm attacking some group using your software (such as another nation) I'm going to flood the fuck out of your PR system till you give up hope and die. I'm going to make you attack your contributors. I'm going to sow confusion so I have the maximum amount of time to lay waste to my enemies and profit to the max.

The internet is hostile. Software is hostile. There are sharks looking to eat you.

Time to face that fact.

by pixl97

5/11/2026 at 4:25:54 PM

> And then there's the team at mozilla

And then there’s the team at curl. Don’t fall for the cheap marketing stuff just because you like them

Everything points to Mythos being marginally better and nobody being able to afford to run it.

by u_fucking_dork

5/11/2026 at 4:47:15 PM

Many can afford it, many companies are easily spending 15-30K a month in tokens per staff.

by addedGone

5/11/2026 at 2:20:39 PM

> Unless you have reasons to not believe the dozens of credentialed, well respected people in the field that have already shared their opinions after working with mythos.

Exactly the same argument was made about o3-preview, lol. But anyway, do they talk about all domains where Mythos did the leap in capabilities (math and other research, ML, SWE) or only about cybersec?

> And then there's the team at mozilla. They wrote a blog about this, and they've worked with anthropic before, using opus 4.6 and found and fixed 22 vulnerabilities. Then they worked with mythos and found and fixed 271 vulnerabilities

Those 22 bugs were found in February, at the time when Mozilla were doing first small-scale experiments with Opus 4.6 (i.e. no proper integration into workflow, likely relatively simple harness, likely only small part of codebase was covered). You can't compare "22 bugs which were found during very early attempts to apply AI" and "271 bugs which were found during large-scale codebase scanning with properly configured AI". The fact that Mozilla is pretty vague about "contribution of other AI models" makes it even worse.

> Unless you're going to accuse them of being shills, these are unquestionable numbers. The model is quantitatively better at this thing

They found another ~150 bugs after their first announce, and only like ~35 were found by Mythos. It's already very sharp drop in contribution.

> I think there are better things to accuse anthropic of, than that they are simply lying for marketing purposes.

Anthropic already used a lot of "technically correct but in fact deceiving" statements in Mythos system card. They are playing both "It's too dangerous" and "We don't have enough compute for that super model" at the moment (it's usually a big red falg). Opus 4.7 (which was likely supposed to be "Opus 5.0", given various facts) is a disaster from various points of views. Of course people don't really believe Anthropic.

by ZrArm

5/11/2026 at 4:54:21 PM

HN readers really like the idea of Daniel, he's sticking it to the man... by doing free publicity and work for the man haha

by doctorpangloss

5/11/2026 at 8:37:00 PM

> curl is one of the most fuzzed and audited C codebases in existence (OSS-Fuzz, Coverity, CodeQL, multiple paid audits). Finding anything in the hot paths (HTTP/1, TLS, URL parsing core) is unlikely.

The way this reads sounds more like the LLM dismissed trying rather than it tried and failed, I've seen Claude do that often unless I probe it to challenge itself, curious here what actually happened.

by hmokiguess

5/11/2026 at 11:58:39 AM

> These tools and the analyses they have done have triggered somewhere between two and three hundred bugfixes merged in curl through-out the recent 8-10 months or so.

If you've just gone through a lengthy analysis of your code with other AI tools, surely it's reasonable not to expect to see hundreds more from a new tool?

It should be possible, unless more bugs are introduced, to eventually get to a state where there are no more bugs in your code.

Process aside, it sounds like Daniel expected to find dozens/hundreds more bugs.

by nevi-me

5/11/2026 at 12:01:11 PM

Mythos was kind of hyped as the tool that would discover much more bugs than any currently available tool

by jaapz

5/11/2026 at 2:03:43 PM

curl had ~15 CVEs in 2026 so far. You surely don't think those (and the one Mythos found) were the last security bugs still left in the code base? There certainly will be more, in fact Daniel predicts ~50 CVEs for the entire year.

But Mythos found 1. After all that hype. 1.

by pbmonster

5/11/2026 at 9:17:58 PM

Maybe curl is just... better hardened? Firefox posted hundreds in April.

by knowaveragejoe

5/12/2026 at 7:25:48 AM

That's not the argument. Yes, curl is insanely hardened. But still, they currently have a new CVE every couple of weeks. Mythos didn't accelerate this much, no more than all the other AI-assisted security analysis they've been doing anyway.

Which either means that, tragically for Mythos, it only got to analyze the code base just after ALL the bugs where finally ironed out and now curl is bug free forever after - or Mythos isn't really all that good, dozens/hundreds more bugs remain and will be found in the next months and years.

I just think the former is a bit unlikely.

by pbmonster

5/11/2026 at 2:33:31 PM

I feel like, if it was a codebase without using any security analysis tools, there would have been some more significant findings - perhaps they can re-run it on an 18 month old commit and see how many it found that were subsequenty found and fixed?

Anyway, I think the case that frontier and next-gen models will get increasingly adept at finding vulnerabilities and that those on the receiving end of those vulnerabilities need to be on top of it.

by tgtweak

5/12/2026 at 5:43:27 PM

Unfortunately that doesn't help much. LLMs are really really good at digging up known vulns, so much so that they often falsely declare known vulns as new and novel ones.

They have the CVEs in their training data, know how to look up ossfuzz logs, etc.

by ostif-derek

5/11/2026 at 1:23:25 PM

If priced like other Anthropic models, Mythos will make vulnerability discovery a lot more accessible.

The author compares it to AISLE, ZeroPath, and OpenAI’s Codex Security. AISLE and ZeroPath are much more expensive. OpenAI’s Codex Security is gated.

Most people don't care about the first two and don't complain about the latter's policy because they are all specialized models and/or harnesses.

Mythos will be available to all.

by andromaton

5/11/2026 at 6:34:29 PM

> AISLE and ZeroPath are much more expensiv

AISLE is *cheaper* for sure

by vibedev999

5/11/2026 at 8:11:12 AM

I don't know about Mythos but in recent weeks I've noticed Opus is constantly failing to fix things in tsz[0] vs GPT 5.5 can easily churn out fixes that are solid and pass tests. I've stopped paying for Claude for now and all my money is going to OpenAI at the moment. Either Opus is massively nerfed or GPT 5.5 is really head and shoulder higher in terms of very difficult tasks. The last percent of conformance tests in tsz are really really difficult and I've seen Opus bailing again and again. So annoying to waste time and tokens to finally get "this is too involved" or "this requires a multi-week sprint to fix".

[0] https://tsz.dev

by mohsen1

5/11/2026 at 8:19:20 AM

The new Opus feels like a step backwards. More expensive, thinks more, and it does not get the job done.

by _pdp_

5/11/2026 at 9:17:01 AM

From a user’s perspective 4.7 is a downgrade compared to 4.6 . It’s intended to give Anthropic more control about their compute resources and profitability:

https://news.ycombinator.com/item?id=48072916

by vincent_s

5/11/2026 at 8:17:32 AM

Having never used Claude and only Codex, does Claude actually say “this is too involved” as a response to a prompt?

by dyauspitr

5/11/2026 at 8:29:43 AM

Yes it does. Usually after hours of working and not getting results

by mohsen1

5/11/2026 at 9:50:29 AM

I am curious, what kind of work do you use Claude for that sometimes requires hours of working. In my case, I have never seen it go off for more than 10 mins and even that is very rare.

by redditor98654

5/11/2026 at 12:03:38 PM

debugging code. I had some issue so I create a plan to root cause that would run the code, change some functions or variables and run again until we get a confirmed answer.

I just work up to that very workflow this morning. I ran last night and finished at around 3am with ~200k tokens spent. Fixed the issue and created a follow up doc for things that it could not verify.

by big_youth

5/11/2026 at 11:45:21 AM

https://github.com/mohsen1/tsz

by mohsen1

5/11/2026 at 4:23:54 PM

Love Daniel's writing style here. Fact based, concise, easy to read

by jorisw

5/11/2026 at 9:26:58 PM

IMHO Mythos was more of a marketing ploy.

When it comes to security and AI, all top tier publicly accessible models (GPT 5.5, Opus 4.7) and even near-top like Deepseek 4 PRO can do a very good job given detailed harness on how to spot issues and cross-validate them to avoid false positives.

by ilia-a

5/11/2026 at 2:16:44 PM

"I signed the contract for getting access, but then nothing happened. Weeks went past and I was told there was a hiccup somewhere and access was delayed.

Eventually, I was instead offered that someone else, who has access to the model, could run a scan and analysis on curl for me using Mythos and send me a report. To me, the distinction isn’t that important."

Really? We're talking about (essentially) a product demo from a trillion dollar industry fueled by debt. Clearly, blog posts like this have an immense influence on the perception of usefulness of the particular model and AI in general. With so much staked on this for the company, wouldn't you want to be sure that you're using the actual product without anyone messing with the results in any way?

by romaniv

5/11/2026 at 1:27:30 PM

I'm disinclined to be overly generous to Antrophic, but I have to say that regardless of whether the talk of Mythos being uniquely dangerous was mostly cynical: It would be great if this starts a trend of giving security-critical software a few months head start with any new significantly improved model.

by Semkas

5/11/2026 at 7:26:12 PM

I guess we miss fundamental information: how much in terms of time and token usage took to the "middle guy" to create the report?

Next question: could it be that OP can use Mythos in a better way since he knows better the project?

by vb-8448

5/12/2026 at 3:08:29 AM

Interesting. curl team found that Mythos is mostly hype while Calif team found Mythos amazing.

I would think Calif (a security firm) is a better team to better utilize such tool.

by tuananh

5/11/2026 at 11:13:57 AM

I'm looking forward to trying Mythos run against my 5000-line, instant-finality, quantum-resistant blockchain project and decentralized exchange (an additional 5000 lines). I already ran all the models up to Opus 4.6 and they couldn't find anything.

by jongjong

5/11/2026 at 8:35:39 AM

I routinely used to compile C programs on other compilers to find defects that one or another didn't find. Compiling on Windows vs Linux. You could summarize / minimize it down to compiling it with warning as errors etc but you'd be missing the point.

The point wasn't actual cross-platform portability even though that was a nice side effect. It was to flush out all the weird edge cases.

Edges like security flaws. Buffer overflows are usually platform specific. There are plenty of other ways to find these issues but simply recompiling for a different platform surfaces all sorts of issues.

by absynth

5/11/2026 at 5:04:40 PM

It's also a convenient way to get press (and investor valuation) for a new model with releasing it (word is they don't have enough hardware to do so).

by tedd4u

5/12/2026 at 12:27:26 PM

Like clockwork, criticism of the Alpha Omega apparatchiks is flagged. They know how to protect their income streams while open source authors get nothing.

by 23aqsI

5/11/2026 at 7:34:34 AM

> The source code consists of 660,000 words, which is 12% more words than the entire English edition of the novel War and Piece.

Typo, or is there a spoof I should go read?

by yjftsjthsd-h

5/11/2026 at 8:17:45 AM

War and Peace is about 590,000 words. Tiny compared to the full Harry Potter collection (about 1 million words over the 7 books), but long for a single book.

by iso1631

5/11/2026 at 8:21:50 AM

They're referring to the typo in the title, "Piece" vs "Peace".

I also thought they were contending the word count before noticing. Even remarked how I find this a weird metric, given that code is not prose [0], but then I deleted that once I picked up on what's going on.

[0] comparing the output of `wc -w` with the word counts of books I'm reasonably sure will be super off

edit: ran a calc, substituting out symbols (but not underscores), digits, and comments yields a 390K word count compared to the 660K cited. not excluding the comments yields 600K, so more than a third of all words in the sources are comments.

by perching_aix

5/11/2026 at 1:55:37 PM

ahh interesting thanks

I guess it's related to the phenomenon where you can read words relatively easily as long as the first and last letters are correct and the rest of the letters are there.

https://wire.insiderfinance.io/the-brains-power-to-read-jumb...

by iso1631

5/11/2026 at 10:01:40 AM

The ten main Malazan books are 3.3 million words, apparently. No wonder it took me such a long time to get through them.

by Accacin

5/11/2026 at 8:01:53 AM

Perhaps he was dictating.

Does it say anything else? Just 'Aaaarggghhhh'?

by dotancohen

5/11/2026 at 8:05:16 AM

Doubt it considering that Daniel Stenberg is Swedish. English dictation when you speak English as a second language with an accent is quite annoying.

by Hamuko

5/11/2026 at 8:33:41 AM

Voice input works really well for people speaking English with a Swedish accent. I think the accent of most educated Swedes is mostly a case of prosody. For sure there are some sounds we say slightly differently than native English speakers. We often have some trouble with /s/ and /z/, but I don't know, "war and peace", I think that's easily understood.

Source: voice typing this with Swedish vocal chords, and only had to correct "different lives" to "differently", and add /[^\w\s]/.

by Tistron

5/11/2026 at 9:09:18 AM

Android voice input works with kids using both English and native words, here in India. The country runs schools in 25+ primary languages, each with dialects, so a TV/phone with voice input is more marvelous than the nitpicks discussed here.

by aitchnyu

5/11/2026 at 9:38:41 AM

I understand completely. You don't want to know what the machine produced, when I asked it for "a new display".

by dotancohen

5/11/2026 at 4:25:24 PM

I personally belive its a marketing stunt and they are just using actual humans to find the bugs/vulns

by plexescor

5/11/2026 at 9:10:09 PM

Who is using Mythos to find these things and where do they run it?

by theaniketmaurya

5/11/2026 at 12:12:27 PM

Swival found many more vulnerabilities without Mythos https://github.com/swival/security-audits

by jedisct1

5/11/2026 at 2:19:53 PM

> (I am purposely leaving out the identity of the individual(s) involved in getting the curl analysis done as it is not the point of this blog post.)

I would very much like to know if they were independent or affiliated to Anthropic.

> My personal conclusion can however not end up with anything else than that the big hype around this model so far was primarily marketing.

... because of this.

by nottorp

5/11/2026 at 3:13:44 PM

AI not finding a security issue on cURL has more to do with lack of widespread security issues than the model's capacity of finding them.

by brunoborges

5/11/2026 at 4:19:58 PM

Should have scanned it with Mythos on an older code base before all these other sec issues was resolved with other tools. Or use the other tools to introduce the same kind of errors in other parts of the code base to see if Mythos would have found it.

A problem is that these tools seems smarter than they are cause they already read seen the answer key.

by AtNightWeCode

5/11/2026 at 3:06:05 PM

Kinda burying the lede: AI tools found over a dozen CVEs in curl last year, and hundreds of bugs.

"Primarily AISLE, Zeropath and OpenAI’s Codex Security have been used to scrutinize the code with AI. These tools and the analyses they have done have triggered somewhere between two and three hundred bugfixes merged in curl through-out the recent 8-10 months or so. A bunch of the findings these AI tools reported were confirmed vulnerabilities and have been published as CVEs. Probably a dozen or more."

by readthenotes1

5/11/2026 at 5:38:30 PM

Not exactly "burying the lede" since Daniel already posted an update about it months ago [1] with extensive discussion in numerous of articles [2] including on this site [3].

[1] https://lists.haxx.se/pipermail/daniel/2025-September/000127...

[2] https://www.theregister.com/software/2025/10/02/curl-project...

[3] https://news.ycombinator.com/item?id=45449348

by toraway

5/11/2026 at 6:47:00 PM

[flagged]

by deferredgrant

5/11/2026 at 1:40:39 PM

[dead]

by sdhrag

5/11/2026 at 8:26:49 AM

[flagged]

by almogodel

5/11/2026 at 9:37:29 AM

Won my bet "voted 10 [vulnerabilities] but in retrospect as you are familiar with Claude and such tooling if you already used any of recent model to done some kind of security review then I'd drop to 1 or even 0." https://mastodon.pirateparty.be/@utopiah/116537456780283420

by utopiah

5/11/2026 at 8:17:19 AM

It's a shame he seems to reject the idea of actually diving in and using these tools interactively:

> It’s not that I would have a lot of time to explore lots of different prompts and doing deep dive adventures anyway.

His expertise I think would elevate the results quite a bit. Although if he never uses LLMs, which it reads like he doesn't, I guess it might backfire just as well. Prompting style (still?) does matter after all, certainly in my experience anyways.

by perching_aix

5/11/2026 at 9:00:03 AM

He states in the article that they use LLMs for this purpose and find them extremely useful.

by jph00

5/11/2026 at 9:26:13 AM

Which can be true without this also being true:

> using these tools interactively

I did read the article. It seems to me they're using LLMs in a prepared manner instead, as mere scanners that produce reports.

by perching_aix

5/11/2026 at 2:14:24 PM

Perhaps I'm misreading something? From my reading of the article, it doesn't sound like Anthropic offered to let him use Mythos in any other way than that.

by SpicyLemonZest

5/11/2026 at 2:21:35 PM

He explains in the article that he failed to actually secure access in the end, even though it was approved. Someone else prompted the model on his behalf, and just passed on the findings.

by perching_aix

5/11/2026 at 10:40:43 AM

He posts about his use of language models a lot on Mastodon[0]. He does lots with language models, but doesn't buy all the way into the hype. I'd say he's one of. most reasonable & balanced voices on the subject of AI use in software today. Happy to use the technology, more than willing to push back on marketing bs.

[0] https://mastodon.social/@bagder

by OtherShrezzing

5/11/2026 at 12:31:47 PM

I see, thanks.

I checked back two weeks worth of posts, reposts, and replies there, and do not see anything suggesting so, so I'll have to take your word for this.

What I do see is him responding to seemingly rather frequent harassment about AI use @ curl however. The stance he takes in those cases is very reasonable (even if you don't use AI for scanning the codebase and contributions, threat actors will), it's unfortunate this topic is so political that he has to deal with this to such an extent.

by perching_aix