8/23/2026 at 7:58:24 PM
The real revolution is Deepseek v4 flash and similar models (GPT 5.6 Luna, muse spark 1.2, mimo, etc...) - Genuinely good performance for a tiny fraction of the cost of Fable and even GLM etc...I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis
by nchmy
8/23/2026 at 9:04:09 PM
I was using ChatGPT voice during cooking to reflect on variations of a dishes i was preparing for years.It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it
by geniium
8/24/2026 at 12:25:03 PM
As long as these models constantly keep switching up things like the temperatures at which a steak will be medium rare or at which temperature to season cast iron, you will never be able to trust them for cooking. My mother ruined a nice waterfowl for Christmas by listening to Gemini.And this is inherent to how LLMs work.
by jorvi
8/24/2026 at 1:37:12 PM
Not if you let your LLM grounds its truth in established facts. Otherwise they would be useless for programming for example.by lukan
8/24/2026 at 2:27:37 PM
I don't understand what this means. I use LLMs daily for my work in programming things, and they regularly will assert things that are not accurate.by saghm
8/24/2026 at 2:55:31 PM
A lot like humans, really. People regularly cite something they read, or quote a stat that turns out to be just completely inaccurate. But if you look up the thing, then you have facts again.by stevage
8/24/2026 at 6:45:30 PM
Sure, and by the same logic, if I was trying to cook a steak, I would not trust an arbitrary human to know the correct temperature off the top of their head; I'd want someone who I could trust had actual experience with the task I was trying to perform. The difference is that most humans are fully able to recognize whether they've cooked steak often enough to know the correct temperature off the top of their head, and they will say "I don't know" to most random arbitrary questions you ask them outside of their experience. I've yet to see an LLM product aimed at general usage for individuals be willing to say this without someone having to literally direct them to give that as an answer if they're not sure.by saghm
8/24/2026 at 2:48:57 PM
When you let the the agent do a test, or tell it to read that doc first, you will ground it in reality. Doesn't mean they are 100% reliable. But without and on their own without access to grounding information, they halluzinate wildly.by lukan
8/24/2026 at 6:42:08 PM
The parent comment said that these types of behaviors are inherent to how LLMs work, and your response was to "ground them in facts". From what I can tell, this does not meaningfully change the original point the parent comment was making that the flaw is inherent; throwing a bunch of extra context at it to try to make it happen less often is useful, but it's still just a best-effort mitigation for the behavior, not somehow a way of literally changing the inherent nature of it.by saghm
8/24/2026 at 2:36:21 PM
So why isn't this the default mode then?by elorant
8/24/2026 at 2:39:38 PM
More expensive.by lukan
8/24/2026 at 12:37:09 PM
While yes, their current reliability is too spiky to be relied upon for a lot of things:Any given failure is not inherent, they are all dependent failures; what is inherent (due to the SOTA in ML, perhaps or perhaps not the architecture) is how many examples they need to get good at stuff.
by ben_w
8/24/2026 at 5:22:10 PM
But how would that be different if she wrecked it following some other online recipe?by kvakerok
8/24/2026 at 2:13:49 PM
Well I wouldn’t trust Gemini with most things.by geraldwhen
8/24/2026 at 2:15:09 PM
Really? I've essentially learned how to cook from Gemini. Not a great cook yet, but I can do the basics now.by ravishi
8/23/2026 at 9:42:30 PM
I share this feeling too. The latest models, even if not necessarily frontier, say Opus 5, Sol high and the likes, I could keep using these models forever even if they did not significantly improve beyond this point. I also believe we'll come up with new ways of using these very same models beyond the mainstream chat and agent interfaces, as the bottleneck is imho in harnesses/environments and not so much model intelligence anymore.+1 regarding voice usage too, I use it in so many different ways it's hard to enumerate: while driving long distances (think of a custom made, interactive podcast) / as a way to collaboratively build specs or shape an idea / as a way to provide input while vibe coding / just as a normal voice assistant (straight in the ChatGPT app or as OpenClaw input via telegram voice notes). I can't overstate how much my routines have changed over the last couple of years.
by adrinavarro
8/24/2026 at 10:24:21 AM
Do you have to give any special instructions to do this? I always want to do something like this, but any time I try I get so sick of listening to what it has to say, just long winded explanations of stuff that tends to go off the rails. Imo it's hard enough to read ai output when I can go back and forward between sentences to make sense of what's being said let alone listen to a continuous train of slop.by drunkboxer
8/24/2026 at 12:12:58 PM
The latest ChatGPT voice mode is really good at being interrupted - I'll often say "no, no, no, that's too much information" while it's talking to stop and redirect it.by simonw
8/24/2026 at 8:25:40 AM
This is why I’m trying to move to open Chinese models — because I will be able to use them forever, while the older Claude models which I genuinely enjoyed writing short stories with have now been deleted, replaced with hypothetically cleverer models which produce text everyone hates.by CJefferson
8/24/2026 at 8:49:22 AM
I also kind of miss how "unhinged" the earlier models were.by olmo23
8/24/2026 at 9:05:22 AM
What about censorship?> I will be able to use them forever
Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
by ShinyLeftPad
8/24/2026 at 10:02:03 AM
Everyone censors for their core jurisdiction/audience. The Enlightened West just calls this guardrailsby rrr_oh_man
8/25/2026 at 4:27:42 AM
"nothing happens in 1989" is censorship. "no I won't tell you how to build a bioweapon for genocide" is guardrails.idk about you but i WANT the second thing, because I like to be alive.
by ShinyLeftPad
8/24/2026 at 11:11:31 AM
You make it sound like the DNC.by euroderf
8/24/2026 at 12:31:39 PM
US models censor and restrict more things than Chinese models by quite a margin.by 59nadir
8/24/2026 at 2:18:41 PM
So you choose bad over worse and pretend it's goodby ShinyLeftPad
8/24/2026 at 11:51:36 AM
> Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?Do you think that fabrication will never progress (in volume) than what we have now? The hyperscalers are already having trouble paying the bills, they can't keep this up forever.
by lelanthran
8/24/2026 at 2:22:14 PM
The hyperscalers will be bailed (maybe not all of them but enough). US economy will crash if not. And whatever is made will go to them, because they pay more (thanks to US taxpayer bucks among other things) than any regular person. First they build on land then they build in space.by ShinyLeftPad
8/24/2026 at 5:10:31 PM
I dislike all censorship, but US models are much more censored, I often find myself using Chinese models to get answers I want.Now, of course I’d prefer no censoring, but I live in the world we live in.
I’m working in the assumption that (like today) there will always be somehow on openrouter, or similar, who will host a model I want to run.
by CJefferson
8/25/2026 at 2:32:29 AM
How do Western models censor?by ShinyLeftPad
8/25/2026 at 2:58:22 AM
I’m guessing they’re referring to things like refusals if it thinks your request may be related to building a bioweapon, or hack someone else, etc.by no-name-here
8/25/2026 at 4:26:24 AM
Some of this seems public safety not censorship.By censorship I mean things like "nothing happens in 1989". By public safety I mean "no I won't tell you how to build a bioweapon for genocide".
by ShinyLeftPad
8/24/2026 at 1:43:33 AM
ChatGPT literally released a major update of their realtime voice model a month or two ago, going from gpt-4o-level (generously) to gpt-5.5 level performance. So at least 2026-level performance was necessary to provide a really good experience.I remember thinking the first ChatGPT realtime voice was science fiction, before the limits on its intelligence (particularly as mainline models advanced) became annoying. Perhaps we’ll feel the same way in a year or two - people have been claiming models are plateauing in practical usefulness every year, and they’ve definitely been wrong so far.
by tmp10423288442
8/23/2026 at 9:12:42 PM
imo this is the problem some of these labs are gonna face, because open models will do this just fine and you as the consumer don't need to pay their training costsespecially considering imo most use falls under this instead of those kind of tasks where you'd need the SOTA
by r_lee
8/23/2026 at 9:37:25 PM
Yeah. Sometimes I wonder who the long term financial winners will be from the ai boom. It might be ram / gpu manufacturers. Or whoever cracks putting LLMs on asics.by josephg
8/24/2026 at 3:12:05 AM
IMO many are still missing a big part of the picture. We're looking at the potential for a massive scale level of automation of [x], which happens to be a huge part of the economy, and people are wondering which player in [x] is going to be the biggest winner. I think the historically precedented answer is none of them.When the Industrial Revolution came along it did create 'super farms' relative to the past through increased efficiency and production, but it also created a huge vacuum in the economy that was ultimately filled by industry, to the point that farming, super or not, became a vanishingly small part of the overall economy - even as production continued to increase.
---
LLMs stand to do the same thing for software. If and when we reach the point of 'normal' people being able to reliably compose ultra customized software solutions to their problems, then software is basically done as a problem-solving industry in and of itself. Not 'done' as in dead, but 'done' as in solved. There's just nowhere to really go from there.
And so I think this will do the exact same thing as the Industrial Revolution did to farming and create a vacuum opening the door to all sorts of new interesting expansions in the real world, as opposed to the digital one. I don't know what this means, because it's quite difficult to foresee the impact of the Industrial Revolution when living in agrarian world, but it's not so hard to see that the future will not be agrarian.
---
So it's probably still myopic but my bet would be on the first major manufacturer of cheap customer/enterprise grade generalized robotics hardware shells.
by somenameforme
8/24/2026 at 8:04:07 AM
I agree that the winners would be doing stuff in the physical world.And there I think the winner would be China.
by petra
8/24/2026 at 8:11:19 AM
We will all be winners.by eru
8/24/2026 at 8:59:13 AM
Maybe - we'll have to wait and see.By my reckoning, there's a significant chance most software engineers will be unemployable within a few years. But I'm not 100% confident that there'll be a utopia waiting for us, as an alternative.
by josephg
8/24/2026 at 9:26:16 AM
Sorry, when I wrote 'all', I meant people all over the world (not just in China).Individuals can still get unlucky. Just like a coal miner might be out of a job, when solar panels become effectively free.
Software engineers are a pretty small part of the general population. And they can move into general white collar work afterwards. Perhaps at a drop in pay compared to software engineering, but still pretty cushy by the standards of ordinary people.
(And if we manage to automate all white collar work to be done cheaply and reliably by machines, well, then we are in utopia.)
by eru
8/24/2026 at 12:55:41 PM
Like all of us we're winners because of the internet?by petra
8/25/2026 at 12:18:59 AM
Similar, yes. Or the industrial revolution.by eru
8/24/2026 at 7:02:51 AM
> We're looking at the potential for a massive scale level of automation of [x], which happens to be a huge part of the economy, and people are wondering which player in [x] is going to be the biggest winn..Sorry to cut you off, but have you looked at Nvidia's numbers since the NFT craze? They won.
Sell shovels in a gold rush, make better shovels, repeat on the next rush.
by x______________
8/24/2026 at 10:43:19 AM
Nvidia has been winning for decades, they got it right in gaming, they got it right in crypto, and they got it right in AI. People (outside of tech mostly) think they just got lucky but if that's the case, they have all the luck in the world.by altmanaltman
8/24/2026 at 12:28:58 PM
Fair competition under capitalism necessarily drives down profit margins; high profits are either temporary, or due to a lack of competition (e.g. someone has a patent or other IP, or regulatory capture). For example, while a lot of the economy depends on electricity: where competition exists, the profit margin for making electricity is not high; where monopolies or government mandates exist, it can be otherwise. This means that assuming anyone wins (i.e. no doom scenario), the winners are probably going to be those who can make best use of the models. Even chip makers will probably not get a long-term boost out of this; there's plenty of room for more efficient compute, and competitive advantages from e.g. ASML last as long as it takes to reinvent their tech, it's not a law of nature.So, my plan would be to invest not in the AI companies, but in the economy as a whole who get to use the AI for their businesses.
Caution though, one thing which AI is already superhuman at is persuasion. Regulatory capture is likely even easier today than one might expect purely from the revenues of the AI companies.
by ben_w
8/24/2026 at 8:10:44 AM
Or perhaps customers / users?Just like Wikipedia put classic encyclopedias out of business, but wasn't really a financially win for anyone.
by eru
8/24/2026 at 9:19:52 AM
I won financially. Encyclopedia sets were expensive.by stbede
8/24/2026 at 9:27:03 AM
Your savings are real, but they don't show up in GDP or a profit-and-loss statement of any company.by eru
8/24/2026 at 10:09:24 AM
The cost of buying the encyclopedia becomes disposable income to be used on other consumer goods. In that sense in shows up in lots of other companies profit-and-loss statements.by KSteffensen
8/24/2026 at 11:04:45 AM
Maybe, but that's very diffuse and hard to attribute to Wikipedia.And it would show up in real GDP, not necessarily in nominal GDP.
by eru
8/23/2026 at 9:38:31 PM
It's going to be the shareholders of the first companies to crack AGI, and make human brains fully irrelevant economically. With the trillions of dollars that's going in through both investment and users, it's going to happen. I don't believe the human brain has fundamental magic that will make this impossible.by a2ff6eeb0
8/24/2026 at 1:56:11 AM
True AGI would upend society in such a way that I'm not sure that being a shareholder of anything would be meaningful. Perhaps being a pitchfork manufacturer is the winning play in this scenario.by adrianN
8/24/2026 at 8:30:19 AM
BRB, longing Remmington and Winchester.by rustcleaner
8/24/2026 at 1:33:27 AM
The true followers (shareholders) of the AI messiah will be saved, everyone else is doomed.by thelastgallon
8/24/2026 at 10:45:01 AM
late stage christianityby altmanaltman
8/24/2026 at 1:09:36 AM
> It's going to be the shareholders of the first companies to crack AGI, and make human brains fully irrelevant economically.What makes you think if one or two AI labs can do this that the rest (including open model providers) won't be able to follow the same path a few weeks/months later?
Even if you believe in the "Singularity", and believe it is coming soon, I still don't see any reason to believe the Singularity will be... singular. There won't be one clear winner, the race doesn't get called as soon as the first person crosses the line.
None of the AI labs are showing any sign of pulling away to a monopoly or duopoly position, to the contrary the early large leads of OpenAI and Anthropic have all been evaporating.
AI has clear economic value. It still isn't clear at all how the providers of AI will capture that value in a moatless environment with the technology becoming rapidly commoditized.
by georgemcbay
8/24/2026 at 5:40:19 PM
Because a month later is a month of recursive self improvement at the speed of light. Once a lab catches up, the first lab will be TWO months ahead, then a year ahead, then forever ahead. After a few months of this the differences in absolute terms will be enormous.by alasdair_
8/24/2026 at 6:21:07 AM
The first AGI that decides it doesn't want any more AGIs is the last one that gets created.by icepush
8/24/2026 at 12:53:03 PM
that would require both AGI and physical bodies for the model, and no kills witches or anything that could stop it, e.g. the militaryeverything would have to be kept under wraps, and you'd need to avoid the scrutiny of the US gov (they already wanna eval SOTA models in advance)
I don't get this idea that "AGI" will just manipulate everyone somehow into destroying the world or something
by r_lee
8/23/2026 at 11:47:00 PM
For the downvoters: What magic do you think the human brain has that makes it impossible to emulate acceptably?by a2ff6eeb0
8/24/2026 at 1:42:37 AM
It's not about the feasibility of the technology.If "human brains become fully irrelevant economically" then that brings into question the entire premise of "share holders" and "financial winners".
What even are money, shares, stocks, and finance in a world where human brains are irrelevant economically? No one knows, but betting that "share holders" will be the winners is a highly questionable bet.
I would much more likely bet that "the armed group who manages to control and benefit from the AI through force" will be the "financial winners" more so than "share holders", who tend to not be terribly military minded at least in America.
by pianopatrick
8/24/2026 at 2:28:30 AM
The AI is likely to control the ability to apply force (see all of the autonomous drone companies). There's a great deal of alignment work being done to ensure that the AI will continue to listen to the shareholders of these companies.If that fails, who knows what things will look like.
by a2ff6eeb0
8/24/2026 at 3:02:20 AM
Are you sure that alignment work is aligning with the share holders and not the operators? Or not the creators? Or not the government? Which of these groups should the AI listen to when these groups disagree?If the AI gets as powerful as you think it might, then the group that figures out the answer to that would have the power, I suppose. or maybe the AI does not listen to any of them and does its own thing. Who knows? Personally, I would not bet the share holders are going to come out "on top" whatever that means.
I think a lot of share holders are finance people, not deeply technical AI people and so odds are the share holders will not really understand the AI enough to be the most likely to control the AI.
by pianopatrick
8/24/2026 at 3:45:41 AM
To be honest: I don't know for certain, but I'd assume that the people who pay the bills get the strongest alignment. They may not be tech people, but I (so far) haven't got a reason to think that the AI engineers are going behind the backs of their corporate leadership and subverting what they're being asked to do; do you?(I think it would be a good thing for humanity if they did)
by a2ff6eeb0
8/24/2026 at 4:56:16 AM
I think right now both the engineers developing AI and the share holders are more focused on beating coding benchmarks and gaining revenue than anything to do with alignment.by pianopatrick
8/24/2026 at 6:35:33 PM
That's alignment with shareholder value, at least.by a2ff6eeb0
8/24/2026 at 3:10:21 AM
The drones don't manufacture themselves, maintain themselves, reload their own ammunition, mine and refine the materials that are used to make them and their ammunition, or operate the power plants needed for all of the above. "AI" isn't going to control diddly squat.by ThrowawayR2
8/24/2026 at 3:47:30 AM
There's a huge amount of research into embodied AI (and, also, people seem to be a lot more ok with manufacturing bullets than pulling triggers).by a2ff6eeb0
8/24/2026 at 7:09:06 AM
You did not read much about history, did you? Or law. If you embed AI into a gun ... you just created a bomb. It is still you who killed whoever it kills.by watwut
8/24/2026 at 5:44:38 PM
Who is going to enforce that law against you, the owner of a massive army of drones?by alasdair_
8/24/2026 at 12:08:23 AM
LLMs are not emulating the human brain. Somebody may well be able to do that someday, but right now nobody is even trying to.by ksenzee
8/24/2026 at 12:09:59 AM
Why would you need brain emulation to get superhuman intelligence?by josephg
8/24/2026 at 12:33:44 AM
Are you making a serious argument that superhuman intelligence is a plausible outcome of training LLMs on everything humanity knows so far? Or are you making the generic assertion that AGI is theoretically possible via means other than emulating the human brain? Because the latter is a strawman (nobody has asserted anything to the contrary), and I have seen no evidence at all to support the former.by ksenzee
8/24/2026 at 12:46:51 AM
I think we can compare the human brain and LLMs on a bunch of capabilities today, and see how we compare. By my reckoning:- LLMs have better long term memory (they know more than any human) and more working memory (LLMs have fast, uniform access to their whole context window).
- LLMs are faster than we are.
- Humans have online learning (we can do simultaneous learning and inference), giving us advantages in many novel tasks.
- We can learn concepts from far less data. And we can manage our mental context more smoothly.
- We seem to have better world models than current models. AI video just doesn't look right, somehow.
I expect that these remaining weaknesses can be overcome without resorting to human brain emulation. I see no reason to think that current LLMs are at the limit of what technology is capable of.
by josephg
8/24/2026 at 10:43:57 AM
Because LLMs don't understand anything. That's the tech. They can only predict what they have been trained with and fail daily at the most basic tasks. Granted they can do amazing things, no question there. But they are not "smart".For example, it seems that even at Fable scale, simple concepts like the passage of time or (gasp) timezones elude them. I live in UTC+10 and with any RFC8339 data LLMs are constantly confused - is it Sunday the 10th or Sunday the 9th, etc. I have tried many solutions for this and every time it finds a way to get it wrong.
by shawnb576
8/24/2026 at 11:18:18 AM
To me it sounds like you're repeating what gp said about the lack of online learning. Do you think that's insurmountble?Getting confused about timezones does not place LLMs behind that many humans. (But doing so repeatedly does highlight the lack of online learning).
by penteract
8/24/2026 at 3:47:12 PM
I see this thread as progress because now there's three people saying this (seemed like it was just me for a year or two).by dboreham
8/24/2026 at 1:38:24 PM
> Because LLMs don't understand anything.How do you square that then? They can do amazing things, but they're also not smart? Do you think its possible to solve Erdos problems without any "smarts"? Can you do it without even understanding mathematics?
I find it very hard to hold the idea that LLMs don't understand anything. They can explain concepts, translate them, simplify them and implement them in code. From the outside, LLMs seem to understands most concepts better than most humans do. Do you understand anything? Couldn't I make the same argument? How would you prove that you understand what a for loop is, or that you know what calculus is? I assume you'd demonstrate your knowledge by using a for loop in a program, or explain calculus back to me. But LLMs can do that too.
> For example, it seems that even at Fable scale, simple concepts like the passage of time or (gasp) timezones elude them.
Funny example, because lots of human struggle with this too. The number of meetings I've had with people in the US! "Lets meet on thursday morning australia time!". Only, they actually meant thursday night US time, which is friday morning australia time. "Oooh that's so weird! Its the next day for you!". ...... Yes, I know.
I think LLMs are just a different kind of intelligence than humans. They're better at some things than us, and worse than others. They can find latent security vulnerabilities in the linux kernel, but struggle to count the Rs in strawberry. They're not as smart as humans in many ways. But we're not as smart as LLMs in plenty of ways too. I didn't find those linux bugs.
by josephg
8/24/2026 at 7:50:11 AM
What is meant by 'superhuman intelligence'? Certainly it seems to be the case that LLMs are capable of a sort of 'polyhuman' intelligence, in that the same LLM that advances mathematics with a novel proof can add unit testing for a new software feature, design a recipe, and create an SVG of a pelican riding a bicycle. As generalists I'd say they're already 'superhuman'.by OJFord
8/24/2026 at 8:16:53 AM
> training LLMs on everything humanity knows so far?That's not all of what we are doing for at least a year, possibly few. LLMs are trained increasingly on generated inputs. Soon human sourced material is going to be rounding error in the process of training.
by scotty79
8/24/2026 at 4:53:20 PM
To clarify: We are not training LLMs on any information that humanity does not already have access to.by ksenzee
8/24/2026 at 2:07:43 AM
Are you making a serious argument that superhuman intelligence is a plausible outcome of training LLMs on everything humanity knows so far?Are you making a serious argument that it's not?
Because you'll need to explain leading-edge mathematics advances that have come from LLMs, among other things.
by CamperBob2
8/24/2026 at 4:58:08 PM
Superhuman intelligence? Really? These leading-edge advances indicate intelligence beyond the level of humanity?by ksenzee
8/24/2026 at 11:56:59 AM
> What magic do you think the human brain has that makes it impossible to emulate acceptably?If I knew, I'd be rich from deploying it onto a substrate for my own AI.
But that doesn't mean that there isn't something there - the current approach seems at odds with how flesh brains work.
I mean, you can power a human brain with 2x bananas for 4 hours, the energy of which might power an H100 for about 20 seconds. It's obvious that there's something different happening.
by lelanthran
8/24/2026 at 11:23:40 AM
this is how i felt about opus 4.6 i still use it it's just faster and does enough to be super helpful. i've used these later anthropic ones a few times but the word salad and slowness feels like it just opens the door to building shit that just stacks and adds on itself.if deepseek and stuff are 4.6 caliber i literally don't know why im here i should probably just go sign up for openrouter at this point
by trueno
8/24/2026 at 4:50:36 PM
the jump from 4.7-4.8 to 5 is so bad in terms of the word saladi just get fatigued from it, am I holding it wrong or something?
sometimes it's fine but the constant RLHFisms like the constant "worth flagging" and stuff is getting really old
by r_lee
8/23/2026 at 11:01:55 PM
All it needs is Internet access to remain useful with few shortcomings.The next step would be automatic self-training. A free LLM that could access HN everyday (and the linked sites) for more data would remain current in programming for a really long time.
by glimshe
8/24/2026 at 8:50:45 PM
This is why it is so important to hoard offline models. They are already extremely capable, moreso than many realize.by tencentshill
8/24/2026 at 12:20:45 PM
That sounds ideal for cooking but I do wonder about programming. Models frozen in amber won’t ever learn new APIs as they become available and development will end up in some weird kind of stasis.by afavour
8/24/2026 at 12:43:26 PM
embers are probably too hot to freeze anythingby PcChip
8/24/2026 at 3:33:39 PM
eventually the novelty wears off and depression kicks inby someothherguyy
8/23/2026 at 8:56:09 PM
>I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster.There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.
by matteoraso
8/24/2026 at 3:41:07 PM
i wish i was experiencing these things that everyone else is.my experience is mostly frustration and rewrites of anything that requires more than what would take me an hour to do myself, unless it is pure translation / boiler plate work.
the leaps are there at getting to more "shaped" code (code that is correct for linters, static checking, etc), but i don't see the models exhibiting much intelligence. i really can't think of a time using LLMs for building anything where they did something that would make me go, "wow, that is really impressive, i wonder how it came up with that." just brute force search and pattern matching still.
even the interesting results in academic work seem to be more of a function of effort (proofs by exhaustion, fitting puzzle pieces in a search space, etc) than anything else. not to say people aren't using large language models to do impressive things, but the agents themselves do not seem very intelligent to me.
it feels like some engineering teams are aware of this fact and are driving agents using strict rule checks (like hooks on steroids), so they can drive some shape of output that aligns with what they require.
by someothherguyy
8/24/2026 at 2:20:09 AM
I have argued for a while that this was an S-curve it was just a case of figuring out which part of it we were in. I am more confident nowadays that we are heading towards the upper plateau but there might still be some head room on that.by ColdStream
8/24/2026 at 2:14:50 PM
I'm really not convinced that these models are even that much more intelligent, as opposed to simply being more token aggressive. I do not find Fable that much smarter than Opus 4.6, and no Opus model seems to have improved things much at all.Benchmarks seem gamed at this point, real world experience just doesn't match up.
by insanitybit
8/24/2026 at 2:46:02 PM
But they don't really have that option. They're trapped in a Red Queen's race.The world keeps moving on, and so the models need to be retrained so that they can keep up with new information. Otherwise you'll get stuck with a model that only works well with information that existed prior to a dataset horizon that's receding into the past at a constant rate.
At the same time, they have to keep iterating on the training process itself. AI generated text and code is slowly spreading across the internet. Model collapse is a real concern; they wouldn't be spending quite so much energy on buying and scanning rare books if it weren't. But for coding in particular expanding their corpus of old text is not really a good option because of the previous problem - no good training your LLM to write 1980 vintage K&R C that won't even compile on a modern compiler.
by bunderbunder
8/24/2026 at 3:20:10 PM
Newer models[1] are being trained in ways that prioritize coding and agentic performance over raw knowledge[2] such that they increasingly rely on external tools for accessing hard data and information.[1] https://artificialanalysis.ai/evaluations/omniscience?models...
[2] https://old.reddit.com/r/LocalLLaMA/comments/1vt7l3e/qwen382...
by cosmojg
8/24/2026 at 3:10:36 PM
Models don't need to keep retraining just to stay current. Harnesses give them access to the internet, internal systems use RAGs, and so on.I think lower-cost models will get the largest piece of the pie, as with almost everything that has ever been sold.
Just look at cars: US consumers buy the F-150, EU consumers buy the freaking Dacia Sandero the most :)))
Ferrari/Lambo numbers are microscopic
by ninahaberl
8/24/2026 at 3:14:33 PM
Even with a harness, models don't reach out for new information they don't know about. For some tech, I have to have a local model draft a plan, then I have to adjust the plan to update it with the new API and references for where to find it. Even if I include that updated information in the prompt for the plan, the model says "what the user says is wrong, they probably meant this instead" and goes off in its own direction with old APIs anyway.by zdragnar
8/24/2026 at 4:12:00 PM
You are basically saying that some models (your local one, which is it?) in some setups (the API you mentioned) can fail to use fresh information if that conflicts with strong training priors. I agree:)BUT
That's a bad model. My opinion is that for exactly this case we need to use RAGs/ APIs/ some retrieval mechanisms.
It's silly to train them on stuff that changes every week/month
I don't learn APIs by heart, I look them up. It's to expensive (my time) for me and (the compute) for the models
by ninahaberl
8/24/2026 at 5:01:34 PM
For starters, it's not just APIs. Like I pointed out with the C example, programming language syntax and semantics also evolve over time.But also, LLMs' use of RAG to keep track of API evolution is limited. You can see this if you watch an agent at work using a well-known library that has a high rate of breaking changes such as Polars or Guava. There's a huge amount of churn on repeatedly writing code that works with an older version of the API and then diagnosing and fixing the resulting compile- or run-time errors. It can burn through quite a lot of tokens, which drives up usage costs.
I agree that, all else being equal, using language model training to bake knowledge that's easy to look up into the system is kind of silly and inefficient. That's actually been one of my top complaints about hawking these LLMs as a sort of general-purpose AI. But the fact of the matter is that's fairly fundamental to how they work, and RAG is arguably just a hack on top of the basic design to paper over this limitation. RAG's limits become pretty easy to see when working in knowledge domains that aren't very publicly accessible, and therefore produce little text that would have been incorporated into the models' training corpora. It can be a bit of a, "Ignore that man behind the curtain!" experience.
And no I'm not just talking about local models. I've seen it happen with recent GPT-5 and Claude Opus series models, too.
by bunderbunder
8/24/2026 at 4:33:37 PM
I most recently experienced this with Qwen 3.8 27b, though I've seen it on several other versions of their local models. It's also heavily biased towards digging into library source code rather than looking at API documentation.To get it to the point of being remotely useful, I've had it start to write condensed fact blurbs into the agents.md file. It doubts itself so much and questions its every decision to the point that it'll literally blow the entire context on thinking alone in anything but the most basic CRUD projects otherwise.
What an earlier generation model would just start doing, it went out to research the source code in multiple libraries just to see if what it was thinking would work... then it said "Hey, I should really just do it" then went back and started researching more anyway, on and on (even on medium thinking level).
If there's a better local model for writing code, I'm all ears.
by zdragnar
8/24/2026 at 3:00:08 PM
> they don't really have that optionI imagine it must somehow be possible to update a model's understanding of recent events without training a completely new model from scratch?
by cj
8/24/2026 at 3:06:02 PM
yeah its called web searchby anthonypasq
8/24/2026 at 6:58:06 AM
I think it really depends - for a lot of things outside of coding and general knowledge tasks even the best models (fable 5 etc.) are not good enough yet: e.g. CAD, PCB design (though getting there on PCB design), ...by dsrtslnd23
8/24/2026 at 8:27:05 AM
A computer that costs $10,000 is impressive. A computer that costs $100 and reaches billions of people changes the world. Maybe AI will follow the same path.by KunYuan
8/24/2026 at 11:54:48 AM
I ask DS4F to make a plan, then check it with grok/fable, build the code, check it with grok/fable, shipby RALaBarge
8/24/2026 at 4:04:04 AM
Current AI is smart enough to help us, but the creators of AK want it to be smart enough to replace us.by jimmydoe
8/23/2026 at 9:14:23 PM
I'd be content if I could get the DS4 flash, luna, mimo level intelligence running on MY low-end hardware completely offline and bearable TPS, not otherwise.by ksh09
8/24/2026 at 1:29:43 AM
It costs about as much as a cheap car to do this well, but it's attainable now, and qwen 3.8 seems to make it possible on a 5090.by ericd
8/24/2026 at 1:22:08 AM
this is the holy grailby nchmy
8/23/2026 at 8:25:53 PM
If they could be cheap+fast and not try to do too much, that's a good spot for me. I don't use the smarter models as much because of cost and because they're still not good enough to let loose on a lot of problems. For assistance I prefer something that can very quickly spit out a specific piece I can review on the spot and keep going. I let smarter models handle things that I treat as external dependencies and don't care how they're written, but in my core domain I'm still mostly hand codingby lilbigdoot
8/23/2026 at 9:08:52 PM
I have a similar process - its just a pair programmer most of the time. I dont understand how people can have a fleet of agents working a bunch of waterfall specs..by nchmy
8/24/2026 at 2:30:37 PM
I more have an agent that I drive to create features, and then a fleet of agents that turn those ad-hoc implementations into refined, integrated code.by ipsod
8/24/2026 at 2:24:08 AM
> content if they never got smarter, and just kept getting even cheaper/faster.I'm definitely not getting smarter. But my tolerance is 1 drink so I'm definitely cheaper. Also as a result, I spend more time training and so I am faster. And yes, I am more content
by intrasight
8/24/2026 at 7:17:56 AM
The real revolution is both. The cost and capability of frontier intelligence will go up AND the cost of "good enough" intelligence will go down.by nbardy
8/23/2026 at 9:11:24 PM
Evidence actually supports that capabilities are leveling off, and cheaper/faster is not really coming. Just log-linearly more capability at smaller parameter counts as they saturate.by poincareball
8/23/2026 at 10:54:13 PM
Please explain why you think cheaper/faster is not coming?All current devices used to run AI are very far from an efficient solution to the problem. What you really want is a pure dataflow architecture, instead of a von Neumann machine. The reason people aren't really making them yet is that when you build one, even if you use SRAM for the weights, you are binding yourself to the dimensions of the model you target -- your chip is only ever going to run variants of that specific model. And SRAM is much more expensive than ROM, so if you want to make a cheap version, you need to design a specific model into silicon.
Once model improvements taper off, the next thing that will happen is everyone will chase speed. There is no physical reason why a mid-sized model could not run at >1 million tokens per second on leading edge silicon, if all computation that can be parallelized, is. No-one will go straight to that, even for a mid-sized model that's like 20 distinct reticle-limited chips. But something like the next version of Taalas HC1 (presumably called HC2?) will probably boost a ~30B parameter model to ten of thousand of tokens+ per second from a single stream within 12 months.
by Tuna-Fish
8/23/2026 at 10:36:47 PM
what do you mean cheaper/faster is not really coming? the cost of the same level of intelligence steadily decreases year over year. computer hardware also advances at the same time enabling cheaper and faster serving (or move to local)by sipjca
8/23/2026 at 9:27:02 PM
not an AI researcher - this is probably true for these "everything" LLMs but I think specialized models are gonna be the next big thingby bad_haircut72
8/23/2026 at 9:51:51 PM
"Specialized models" are a bit of a doozy.The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very little performance, and for too much effort.
Generality and intelligence seem to be entangled very heavily in LLMs.
by ACCount37
8/23/2026 at 10:18:21 PM
And yet, there's VibeThinker 3B to bring this long-held premise into question (if not to blast it to pieces.) It is practically illiterate by the standards of larger models, yet performs like models 100x its size on mathematical and logical reasoning tasks.by CamperBob2
8/23/2026 at 11:35:19 PM
Which are the kinds of tasks computers have been historically quite good at.It's impressive that it does what it does, don't get me wrong. But if you expect it to replace the likes of GPT 5.6 Luna, let alone Sol? Nah.
by ACCount37
8/23/2026 at 11:48:15 PM
Computers have historically been good at answering word problems fed to them verbatim?by CamperBob2
8/24/2026 at 8:06:29 AM
No, but they were good at answering formalized versions of the same word problems.What this tells us is that a 3B LLM can retain enough NLU to understand those word problems. Which isn't particularly surprising?
And also that the same LLM can solve a math or logic problem it understands. Which is a lot more impressive, because early LLMs were already quite good at NLU, but notoriously bad at things like math, logic and iterative problem solving. This 3B model existing tells us we're beginning to figure out how to imbue models with those capabilities reliably.
by ACCount37
8/24/2026 at 4:40:20 PM
"Formalizing the problem" is pretty much the whole shooting match. It takes intelligence to do that. The rest is mere calculation.by CamperBob2
8/24/2026 at 8:25:07 AM
Computers were never good at math. They were good at pre-coded algebra.When LLMs started to get popular, they really were stochastic parrots. I was fully aware that they were completely useless (except perhaps for poets) until they can do math. And I was a bit skeptical that they will ever be able to do math. But they started to do math and recently they got really good at it.
Math is the pinnacle of human achievement. You can't do anything harder with your intelligence than math. And LLMs are now doing it.
The fact that 3B model is capable of doing math on the level that is better than what frontier models trained for millions could do 3 years ago is absolutely stunning.
by scotty79
8/24/2026 at 9:15:55 AM
Moravec's paradox begs to differ. Things that are hard to humans are easy. Things that are easy to humans are hard.Math is incredibly hard to humans, but "proving a conjecture" might have a lower intrinsic complexity than "putting together a good joke". It's just that evolution has only ever optimized for one of those things.
Math can easily end up being one of those things that are less "hard" than they are "hard if you're a meat-brained hairless ape" - like chess play did.
Historically? "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division. Then it suddenly didn't require any intelligence whatsoever.
Don't get me wrong - the LLM achievements in math, both as in "solving unformalized problems" like VibeThinker does and in "rolling novel math" like the latest ChatGPT and Fable do are very impressive. We're come a very long way from "formal logic only" systems of the 90s. The AI progress we see now never ceases to impress me.
But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing - so be wary of your intuition when saying things like "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
by ACCount37
8/24/2026 at 2:34:00 PM
> but "proving a conjecture" might have a lower intrinsic complexity than "putting together a good joke"Might earwax be soon worth more than gold? Experts say: No! What? No.
> "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division.
No. Not intelligence. Diligence. https://en.wikipedia.org/wiki/Computer_(occupation) They didn't hire the smartest to do the calculations. They hired the diligent and cheap. They hired the smartest to do math.
Things that are easy for humans are easy because we have fist sized universal approximator in our skulls, that's just fast enough to keep most of us on two feet, architecturally optimized for very few activities (mostly physical, some virtualized) and trained for years. It doesn't mean things we do are complex.
As for Moravec's paradox ... Guidance system of a missile is not super smart or solving complex problems. It's just brutally optimized for the task and has a fitting form factor. Tasks that are easy for it are hard or impossible for my windows computer and vice versa. Paradox comes from stupidly thinking easy<->hard is one dimensional axis. That kind of thinking is something people are very prone to ... good<->evil, healthy<->sick, young<->old ... while if we go a bit beyond the simplest narratives we can plainly see that everything is a multidimensional landscape. Just because we chose to draw a single line through it, in a semi-random direction we feel is about right, doesn't mean it is relevant for solving anything or even interesting. That's where a lot of paradoxes come from. We just strayed from reality too far and simplified or abstracted something too much.
> But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing -
I think we can get a good hang of estimating how hard a thing is. If a thing is hard for a human it's probably pretty hard. We made some of them easy building machines that exceeded human strength and diligence. Now we built first one that exceed human intelligence. On one hand, it's as big as invention of a lever, steam machine or a computer. On the other hand it might be only roughly as important as those things.
... If a thing is easy for human it still might be hard because of hardware optimizations that humans have. Walking on two legs, seems easy. Walking on two arms. Much harder. But truly they are one and the same thing for a robot. So you might easily estimate that walking is not that easy. It's just when it comes to legs humans have a specialized controller, like a missile guidance system. Putting together a good joke? Might seem easy, maybe it's not that easy because humor plays a role in reproductions so we might have some optimization for it, but it's surely not harder than putting together quantum theory. You can see this from whatever the ideas version of cyclomatic complexity is. Some math theories have higher complexity than quantum theory. So a system that's capable of exploring multidimensional landscape of mathematic language, surely has raw capability of doing everything else humans can do with language. And it will once we direct it towards it correctly.
> "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
I don't agree it has. And I stand by it.
by scotty79
8/24/2026 at 2:19:28 PM
> Evidence actually supports that capabilities are leveling offWhat evidence?
by naasking
8/23/2026 at 10:38:42 PM
Cheaper/faster is coming for sure.Model on a custom silicon: https://chatjimmy.ai/
1-bit models that run on a CPU: https://github.com/microsoft/BitNet
by ForHackernews
8/24/2026 at 7:57:45 AM
Blazing fast...but terrible. Put Sol on silicon but will still need access to the internet...so it will be somewhat slow anywayby saturn8601
8/24/2026 at 11:10:05 AM
It's terrible because it's Llama 3.1 8B. It's such a crappy model because HC1 was a relatively low budget proof of concept.The team that built is working on a better implementation.
by Tuna-Fish
8/24/2026 at 6:24:49 PM
Not sure its that to be honest. It seems like maybe its not installed correctly or is like GPT-1/GPT-2 quality? I asked it who is [famous actress] and it started talking about some random person from Mexico with a completely different name. The speed is intoxicating but i'd like for it to actually answer based on what I asked. Thats why I think something might be wrong in implementation on this site.Edit: I went back and retested it. It revealed that its data is from July 2021 which explains partially why It couldn't talk about the actress I asked about (she exploded in popularity in 2026 but was still a professional actress in 2021 so idk). I then went and asked questions about a very popular actress and movie in 2010. It got it much better but still hallucinated a ton of details about her.
I guess I didn't fully understand what you were saying. Sorry about that! I look forward to their next releases because upon thinking about what I experienced here, I am super excited to see this progress further!
by saturn8601
8/24/2026 at 10:06:48 PM
It has no internet access and only 8B 3-bit parameters. That's simply not enough for it to compress all that much knowledge.by Tuna-Fish
8/23/2026 at 9:44:34 PM
What "evidence"? Because we keep running out of benchmarks to distinguish frontier model performance. If capabilities are "leveling off", we're not seeing it yet.by ACCount37
8/23/2026 at 9:51:54 PM
Eh. I don't think Luna is good enough. I think that threshold is around Opus / Sol where it can do most of the tasks for me. But I still have many tasks which require either better intelligence or better UI design capabilities.With how generous subscriptions are, what I actually want is GPT Astra, not cheaper Sol.
by redox99