7/25/2026 at 10:19:20 AM
The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model.A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task.
For e.g. you might use Fable for UI design, Sol for systems design backend work and Kimi K3 for exploit development.
The only purpose these metrics serve is bragging rights for the model companies.
by meander_water
7/25/2026 at 11:31:40 AM
A year ago a new Best Model came out, and I was very excited. I tested it on a simple programming task (which required making 3 trivial changes in 3 files).The model did fine.
Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less.
In this moment, andai was enlightened.
by andai
7/26/2026 at 7:56:29 PM
I agree and this is why I think open models will win in the end. There is just so much to gain on being 10% behind the curve. Especially when the curve is far beyond your needs.by jug
7/25/2026 at 2:49:59 PM
Sure, that's true until you hit a difficult problem where the smaller models thrash endlessly whereas its big sibling can solve it with one prompt.by the-grump
7/25/2026 at 1:38:43 PM
is there a benchmark that uses prices or speed as one of the axis, in addition to accuracy?Best could mean different things to different people.
by mejutoco
7/26/2026 at 9:58:23 AM
http://deepswe.datacurve.ai/ has graphs for Cost, for Token Usage, and for Agent Steps. They don't have one for Time, sadly, but Output Tokens and Agent Steps (which appear to produce near-identical rankings) are a decent proxy.https://cognition.com/blog/frontier-code has a dropdown selector for Tokens, Cost, Time, Agent Steps, and more.
---
As a side note, these two benchmarks appear to be more sensitive at distinguishing supposedly frontier models from each other. But they themselves cannot agree on which is better!
One argues that the other has a flawed methodology (and makes a fair case). However it might also just be that the frontier is a little jagged, and sometimes one model will do better than another.
That's been my experience anyway. When I have a very important task, I always make sure to get a second opinion: I get them both to give it a shot, and then to critique each other's solutions. You can actually apply this at every stage (the review, the planning, the implementation), if you have the patience for it.
(Would be nice if there were a way to automate this process. Maybe with one of the higher level meta-harnesses that runs Claude and Codex as subprocesses. But I haven't looked into that yet...)
At any rate, combinations of models have always been found to do significantly better than a single one (e.g. see Model Alloys, Model Fusion, etc.). So make sure to use them, when appropriate!
Hope this helps.
by andai
7/25/2026 at 2:33:07 PM
Yes same website https://artificialanalysis.ai/models#intelligence-comparison... but they don't have graphs for the individual benchmarks sadly.by artemisart
7/25/2026 at 12:39:43 PM
Yes. I do wish there were benchmarks for specific tech stacks. I.e, if I have an Elixir/Phoenix project, which model performs best (idiomatic, etc.) in 2026?Of course it will be somewhat subjective. And I can hang around those communities for opinions. But it might be useful in a world where it's impractical to constantly compare them all, and it varies pretty widely.
by kylecazar
7/25/2026 at 1:07:14 PM
This absolutely should be a thing, but it'll have to be a per-community thing, them building their own dataset and creating their own evals (similar to how people do for production workloads).Although:
"(idiomatic, etc.)" I don't think that should be part of the aim (or it should be under-weighted), because... you can just provide guidance on how to do things more idiomatically, rather than depend on that knowledge already being encoded in the model. I'd be more curious about verifying that it can work through gnarly bugs / features in an Elixir codebase. After all, what use is a model that by default does everything idiomatically if it can't figure out some small concurrency bug.
by pocketarc
7/25/2026 at 7:21:15 PM
It should become a win/win mechanism where communities are rewarded for high quality, human driven validation of LLMs, and vendors gain for bragging rights about their models being highly-skilled-human-approved. Still don't know why this isn't becoming a thing.by 0xCAP
7/25/2026 at 3:37:25 PM
"completely meaningless" "the only purpose"There's some kernel of truth to what you are saying, but hyperboles like this just aren't accurate. All statistics lie but its better than being blind... What your post really says is that benchmarks only show an average over multiple tasks. Yes, obviously, the point of a statistic is to summarize.
by hellohello2
7/25/2026 at 4:13:28 PM
>All statistics lie but its better than being blinddisagree
by dominotw
7/25/2026 at 1:30:55 PM
If you click on the link you will see that it's not "one single metric" there is literally all the metrics so you can make an informed decisionby dbbk
7/25/2026 at 10:33:56 PM
How dare you sir. You can't expect reason and due diligence when it's easier to just be angryby halJordan
7/25/2026 at 11:28:22 AM
Theres value in some of AAs charts, like cost per job, and how often it hallucinated..But I agree, wrapping that up into a single result.. you lose all the nuance, it's just bragging rights.
by intothemild
7/25/2026 at 1:59:59 PM
This sort of rhetoric appears for every benchmark, and somehow it always rises to the top. A few days ago there was a Geekbench 7 submission on here (https://news.ycombinator.com/item?id=49025812), and again the top comment was someone dismissing it, using the classic "but I want a benchmark specifically for exactly the thing I do" perfect-is-the-enemy-of-good nonsense.I, one of those end users, absolutely use these benchmarks as heavy input considerations. Indeed, the vast majority of people do. "Completely meaningless" is just nonsense, of course, and while it doesn't perfectly map to every use, there is a pretty good correlation with suitability for specific tasks.
I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.
Not to mention that the linked page includes a pretty broad list of specialization benchmarks.
by llm_nerd
7/25/2026 at 2:09:21 PM
Firstly, I don't have many issues with benchmarks per se. But I do have issues with leaderboards. And the AA index is touted by lots of people to argue that X model is better than Y, which I find inaccurate.> I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.
This is kind of my point. The benchmarks say they are splitting distance, but they actually vary wildly in performance for specific tasks, so they are in fact not equivalent.
by meander_water
7/25/2026 at 2:27:52 PM
>This is kind of my point.It's a shit point, then. And absolutely no one said they were "equivalent", and again you're doing the rhetorical "it isn't perfect and absolutely comprehensive for every possible scenario, therefore it is "completely meaningless". Again, you chose that absurd terminology, rather than for instance "doesn't tell the whole story".
Again, you chose three models for your example at the very tops of the leaderboards. The SOTA models. Which kind of means that the leaderboards actually mean an incredible amount, no?
by llm_nerd