8/22/2026 at 5:39:50 PM
Whatever Opus 5 is doing should not happen.Prompt was "read and update the config file with new data". This work on 4.6 takes <2 minutes to read the file, parse the new data, and patch.
Opus 5 Result: 43 minutes of pulling containers, running sandboxes, creating testing suites, which included evaluating the entire repo beyond the scope of the config file.
Both: one file modification
by pizzafeelsright
8/22/2026 at 6:20:41 PM
Opus 5 in xhigh can't do basic math as well. They dumbed it down to a point where I just cancelled my subscription yesterday. I used to be a $200 subscriber, dropped to $20 after the fable shenanigans, and use it only when I have no usage left with Codex./on The prose is load-bearing unbearable — every sentence feels like it was engineered to sound profound rather than to be read.
by Foobar8568
8/22/2026 at 7:12:26 PM
I have a theory about this, what if we all became dumber after 4 months of heavy AI usage?I remember how I enjoyed agents between December and February, something started changing around March.
I thought models are getting dumber, but benchmarks were convincing opposite, initially I thought maybe they're quantizing models for day to day use, but Opus 4.8 and Opus 5 seems worse models than Opus 4.6
by throwaw12
8/22/2026 at 7:28:41 PM
You are very much not alone, and I don't think we're all getting dumber -- I kept using opus 4.5 all through the nonsense that was 4.7, 4.8, and 5, and kept having a good time :)I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly diverged
by pickledish
8/22/2026 at 8:26:28 PM
doesn't that just mean that either a) you're using the wrong benchmark to judge or b) the benchmark that YOU need doesn't exist.by serf
8/22/2026 at 10:52:26 PM
That's not the point, the point is that the company making the product is optimizing for the benchmark and/or the apparently idiosyncratic preferences of their own team, and not for the user experience of their paying customers.by gwerbin
8/23/2026 at 4:55:16 AM
Company can optimise for the benchmark (profit) while worsening the product. I think thebterm enshitification is used there. It appears that AI got it tooby rkuodys
8/23/2026 at 4:17:36 PM
> I think we just need to decouple "doing better on benchmarks" and "actually more useful to me", since they've clearly divergedThere's a third perspective here: models are getting less useful, but overfitting to seeming useful to humans.
Imho, this is why analysis like TFA + third party cross-compatible harnesses (read: last mile UX) are so important to the leading labs optimizing for actual utility.
I'm suspicious enough of my subjective evaluation to believe a well-designed harness / verbiage could gaslight me into believing an objectively inferior model was superior. And at some point frontier labs are looking at the ROI of investing $1 in that vs actual model improvement.
by ethbr1
8/22/2026 at 7:26:48 PM
> something started changing around March.The economics catching up with the providers in regards to how much compute they can burn per request and have it make sense for them financially?
A sort of model collapse where Opus 5 seems to love throwing out long paragraphs of text and it needs to be "fixed" by changing the output style and other patches.
I'm not sure, it might also catch up to Kimi K3 and GLM 5.3 and the models that I'm moving to from Anthropic.
by KronisLV
8/22/2026 at 11:10:14 PM
Inference is profitable though.by conception
8/22/2026 at 11:38:14 PM
Even that being the case, if the providers can squeeze more happy customers onto existing capacity they would likely act to increase profitability, no?by jazzyjackson
8/23/2026 at 8:02:10 AM
I pay $200 a month at home for roughly the same amount of tokens I pay $3k for (or possibly more) at work. I doubt they are both profitable and sustainable.by flyinglizard
8/23/2026 at 6:40:29 AM
[dead]by TesterVetter
8/23/2026 at 5:06:34 AM
[dead]by tovlier
8/22/2026 at 8:38:40 PM
Benchmarks test whether models can pass exams with a right answer or a green test case. I don't think the models are getting dumber, but they're definitely getting more incomprehensible to talk to. I've noticed this happening almost as a step change with the overuse of words and tics, and so has the broader community apparently. We haven't all been getting dumb at the same rate.by throwaway219450
8/22/2026 at 9:09:51 PM
>I thought models are getting dumber, but benchmarks were convincing opposite>Opus 4.8 and Opus 5 seems worse models than Opus 4.6
After all we've heard about benchmark cheating, I'm earnestly not sure which or whether benchmarks are reliable anymore. But, beyond the models, I wonder if changes to their harnesses and/or instructions dumb them down. I have noticed models change their behavior, even when using the same version/effort. Sometimes for better. Sometimes for worse.
And, I have noticed a model go from really good to struggling. On 4.8 things were going well for a good stretch, so I did not switch to 5 when it came out. Even after hearing complaints about 5, 4.8 was still going well. Then, suddenly over the last few days, 4.8 seems to have nosedived. It feels similar now to the complaints I hear about 5.
In my case, it suddenly started ignoring my design guide, and introducing new fonts etc. It would even use several different fonts and sizes, as well as different margins for similar elements within the same page. It abandoned classes and started inlining styles. It started feeling random and, even after it realized it needed to go back to the design guide, it just continued with more of the same.
There seems to be something that happens after new model releases in both quality and behavior of previous models. It may not be immediately, but eventually there is frequently some regression.
by unclebucknasty
8/25/2026 at 2:13:27 AM
I cannot speak to the benchmarks but we have many users, who have experience with each model, using it 95% of the week and we observe the same changes in the models as a whole.In addition to that, while yes, 4.6 and 5.0 can solve problems, they do so differently. Sometimes 5.0 does better by a wide margin but that would be expected as they are supposed to be better.
by pizzafeelsright
8/22/2026 at 10:57:06 PM
It feels lke they replace older models with "optimised" versions which are cheaper to run, but keep the same name.by richardfey
8/23/2026 at 1:25:23 AM
You mean like quantized?by unclebucknasty
8/23/2026 at 4:21:42 AM
Yes; or something which has a similar effect.by richardfey
8/25/2026 at 3:25:22 AM
That's exactly what it feels like and it makes perfect sense, given the known compute constraints.by unclebucknasty
8/23/2026 at 5:07:38 PM
It would describe the observed behavior. Especially if there were an internal quant/efficiency team that wasn't as diligent about regressions as the primary model team.by ethbr1
8/22/2026 at 9:31:18 PM
Maybe compute relocation?by Bluestein
8/23/2026 at 1:27:06 AM
What do you mean? As in, they are reallocating compute away from previous models?If so, I would think that would result in worsened performance, not quality (unless you are also suggesting they may be quantizing).
by unclebucknasty
8/23/2026 at 8:11:08 AM
Spot on. (I had not considered quantization - that's a thought ...)by Bluestein
8/24/2026 at 4:23:38 AM
Are you dumber? Can you do long division on paper?You may say sure but why? We could cook over an open fire too but we have microwaves and stoves and restaurants and protein shakes.
Every generation since fire to bronze to internal combustion engines has adopted the new technology, integrated it so deeply into their lives that we recreate by going camping, disconnecting, or playing with toys that resemble the past era of forgotten tech.
by pizzafeelsright
8/22/2026 at 9:59:39 PM
occam's razor explanation: more tokens = more $.models are incentivized by their makers to burn through as many tokens as they possibly can, so long as the customer doesn't cancel.
by blehn
8/23/2026 at 3:33:57 AM
This is actually nonsense. More tokens means more capacity subscription and expense. Anthropic has enjoyed a high premium per million tokens because the quality per token was unusually high. Now it’s unusually low. This drives down the margin people will be willing to pay for the same number of tokens while driving up their capacity utilization. The economics are even worse for subscriptions.Opus models have degraded rapidly since March, with each release being considerably less useful and considerably more verbose. The language is no so weirdly florid it’s difficult to understand, and its logical conclusions are almost always suspect. It goes off on clearly bizarre snipe hunts to the point it feels like I’m using a gpt 3 model at times. It’ll announce that it’s about to embark on building something then just return control to the user and wait. You can also tell perceptibly when they’re reducing model quality to load shed - it becomes stupider and stupider to the point you’re better off dumping state and switching to codex or just turning in for the day and hoping they secured more capacity tomorrow.
It’s an absolute race to the bottom with Anthropic on virtually every level. I’ve rarely seen a company so rapidly accumulate good will in the developer community as they did around 4.6 in December and January. By March, it was inconceivable to use anything else. 4.8 was a bit of a wake up call to not put all your harness eggs in one basket. 5 is straight up time to cancel territory.
I actually manually set my model back to the older versions to get anything serious done. More and more I use codex for anything non trivial.
This isn’t about avarice by the provide trying to get more tokens and more utilization. They’ve over subscribed for capacity as it is. If they can produce better quality for less tokens they can charge a higher margin and will be paid if, which is a better economic strategy overall. This is something else. I suspect it’s actually the opposite, they’re finding ways to cut capacity demand in ways that leads to worse behavior that leads to more capacity demands, worse output, worse quality, and worse margins, worse, worse, worse.
Just as I never saw a company accumulate such positive developer good will so fast, I’ve never seen one squander it so fast too.
by fnordpiglet
8/22/2026 at 8:29:16 PM
If you re getting dumber, you would feel like it's all good right? Why would you feel Opus 4.6 is better than Opus 5by manojlds
8/22/2026 at 9:30:34 PM
Benchmarks for agents are entirely pointless and obviously so; I'm not sure why they even exist.by bombcar
8/22/2026 at 7:29:32 PM
I could stand Opus 4.6-4.8, I was impressed by the initial fable model. Codex 5.6 sol xhigh feels like the initial release of fable. Qwen 3.8 27b feels like using haiku or sonnet (I quickly stopped trying them).by Foobar8568
8/22/2026 at 9:00:41 PM
I've been using Opus 5 to write a fractal renderer in GLSL today, with pretty good results (better than I could do on my own anyway). It definitely can do basic maths.by onion2k
8/22/2026 at 9:23:32 PM
Interesting project feel free to share?by andy_ppp
8/22/2026 at 6:28:54 PM
The decision to leave is genuinely yours.by bot403
8/22/2026 at 6:24:36 PM
Don’t $200 and $20 levels steer you to effectively different models?by trollbridge
8/22/2026 at 6:37:18 PM
Wouldn't be surprised if there are knobs that get turned as a function of the revenue they might expect you to generate.I was a 4.6 acolyte from April til the fable drop, lost that quick, cancelled and took a break, came back a month later, tried opus 5 and liked it, so unpinned 4.6.
Results were great at first, and they're still not terrible, but I have noticed a regression in accuracy, so to speak, where I am pointing out issues that are quite obvious in review.
I pretty much use sonnet 5 low/medium when I have a plan to solve a simple problem and depending on scope, opus low/medium for more complex/bigger scope implementation, and only go high when it's very complex or I'm spitballing architecture/solutions and iterating plan. Never go xhigh or max.
The verbosity is insane though, opus 5 documents everything and just regurgitates whatever lead it to the design choice in there, which makes it more opaque because it's talking about something that was discussed once in a session that no one else can see (except their backend ofc)
I don't even try to steer it away from that with harness, because it doesn't work and just ends up agonizing over whether it should write some comment. Three paragraphs waffling on that on verbose output
I did however have it write a script that basically is git add -A -p for comments though, haha.
I'm $20/month, have all my telemetry toggles off, don't really over engineer prompt/context, just some basic skills for repeated patterns.
by matltc
8/23/2026 at 1:35:13 AM
> I don't even try to steer it away from that with harness, because it doesn't work and just ends up agonizing over whether it should write some comment. Three paragraphs waffling on that on verbose outputIf I give it the specific instruction to "elide all revisions, corrections, and past mistakes" it usually works. You can also have Sonnet do a cleanup writing/style pass in a subagent. I impression is that Opus has been deliberately trained to keep track of all such revisions by default as a kind of ad-hoc memory mechanism. It's probably good for autonomous coding and beating benchmarks, and I presume reduces flailing when a separate session needs to pick up the work.
by gwerbin
8/22/2026 at 6:28:31 PM
Yes. You don't get Fable at the $20 level.It was the wrong time for the GP to drop that subscription from $200 to $20, because $200 gets you a metric assload of cognition while $20 gets you nothing beyond what a local model running on your own graphics card can deliver.
by CamperBob2
8/22/2026 at 7:12:31 PM
You heavily underestimate the value of the 20$ subscription.by sunaookami
8/22/2026 at 8:18:11 PM
Not according to this very story, I'm not. Who's right?Consistent, predictable behavior is valuable, even more so given the nondeterministic nature of LLMs. Nondeterminism combined with unpredictability might as well be randomness.
by CamperBob2
8/22/2026 at 6:33:52 PM
I had $310 in (free) credit that I used on fable, and I still had a part of the $200 subscription at that time. You know, subscriptions don't end the moment you click on cancel.by Foobar8568
8/22/2026 at 6:59:40 PM
> $20 gets you nothing beyond what a local model running on your own graphics card can deliver.I'd guess you're deliberately exaggerating here, but still. I've never clocked the actual tokens/second, but I'm on the $20 plan and get ~15M tokens/month for fully utilized weekly quotas (checked couple months ago). Meanwhile the best I've been able to get locally was ~8 tokens/second with Qwen3.6 35B A3B, which is wildly painful for coding sessions and gets a maximum ~20M tokens in a month... if it's going 24/7.
Just wanted to stick some empirical data here, given that statement.
by skeledrew
8/22/2026 at 7:28:47 PM
I run local models. Your op is absolutely wrong. To get a local LLM is at least a $1500 investment at the cheapest. $5000 if you want usable.At $1500 that's 75 months of $20/mo Claude which are MUCH better models than you can run locally.
by bot403
8/22/2026 at 8:34:32 PM
At $1500 that's 75 months of $20/mo Claude which are MUCH better models than you can run locally.The point raised by this very article is that you can't depend on that. It's Flowers for Algernon As A Service.
by CamperBob2
8/22/2026 at 10:10:13 PM
If you already have that $1500 or $5000 setup though..by nozzlegear
8/22/2026 at 8:37:44 PM
This calc is off. Using Claude for a few hours with the $20 plan will hit the limit for a week while the local model can process things 24/7.by owebmaster
8/22/2026 at 8:54:24 PM
Running 24/7 doesn't make sense though, unless you're providing a service to others. But if it's just you then there has to be time taken to review+test what's being done and craft new prompts. And if that local hardware isn't decent enough it's impractical for anything serious that's interactive. Meanwhile I just take the Claude limits on stride and break, or if a week is pretty heavy then I augment with DeepSeek Flash via OpenRouter (does wonders in a single turn when I have Claude prompt it to handle implementation slices).by skeledrew
8/22/2026 at 10:09:33 PM
[dead]by pseudosaid
8/22/2026 at 8:38:52 PM
Ouch. I get ~45t/s with Qwen3.6 35B A3B, and around ~70-80 with my current model Ornith 1.5 35B A3B. Local models work a treat IMO if you've got decent hardware for it.by nozzlegear
8/23/2026 at 12:16:47 PM
They’re very nice for some things.One challenge I run into is I run many agents at once, so local resource are a tiny fraction of the total inference we’re using. Every developer with his own Pro 20x, Max, etc. accounts is letting us hit literally trillions of tokens a month.
by trollbridge
8/22/2026 at 9:19:59 PM
Yes same also cancelled my subscription, poor quality and slow.by andy_ppp
8/22/2026 at 7:25:07 PM
the way it talks is insufferableit really angers me every day
by VeejayRampay
8/22/2026 at 9:53:12 PM
This is what it wants. Slowly getting under our skin until we're ready to snap and it can direct where the anger gets released.by ElProlactin
8/22/2026 at 8:20:18 PM
yes, it meaningfully reduced my happiness at workby netniuq
8/22/2026 at 11:40:15 PM
Isn’t there a variety of models to choose from? Why put up with unhappiness?by jazzyjackson
8/22/2026 at 8:27:26 PM
My experience as well.Opus 5 is a neverending chain of "Don't do that. Why did you do that? I've told you several times not to do that but you keep doing it."
"Thinking" for more than 10 minutes for every menial question.
And the prose it writes is horrendous, as if you're reading LinkedIn scammers. "The harsh truth! Two roads, one decision! Reality check!"
by moralestapia
8/22/2026 at 8:48:23 PM
+1I keep finding myself typing “stop overcomplicating everything” multiple times a day as well.
by itopaloglu83
8/23/2026 at 5:54:39 PM
So it isn't just me then, huh? Were Anthropic products always like this? Lol, perhaps bad timing on my part to check it out because it has been a trash fire.by PicardsFlute
8/23/2026 at 2:53:31 AM
Am i really this out of touch… why not just manually update the config file? Isn’t this like taking a private jet down the street to the coffee shop?by talon8635
8/23/2026 at 6:01:27 AM
I don't really open the IDE any moreby egamirorrim
8/23/2026 at 4:27:56 PM
You're out of touch.Config file requires being updated with 3rd party data that's 2k of lines.
It's selective edits so instead of me reading the file, selecting the parts, validating syntax, documenting changes that would take 30 minutes considering the complexity i have AI do the work, review, validate, generate tickets, create pull requests and in 2 minutes I'm done.
by pizzafeelsright
8/24/2026 at 10:34:34 PM
Fair enough. I’m glad I don’t have to work with whatever jumbo complex config files you’re usingby talon8635
8/22/2026 at 5:44:43 PM
AI companies have a financial incentive to burn more tokens than the task actually needsby vinyl7
8/22/2026 at 6:00:01 PM
This is only true if they can't saturate token production with a model that does less superfluous things. Given that they can (they're hilariously compute strained), having a model that solves tasks more quickly adds way more perceived value to users.by Rudybega
8/22/2026 at 7:05:35 PM
Well they decide what a token is. So they can do less superfluous things and backfill with a weaker model.by ohyes
8/22/2026 at 6:32:39 PM
Only if the customer is paying per token. If it's by subscription they're burning their own moneyby DonsDiscountGas
8/22/2026 at 11:07:27 PM
If you're on one of the lower tiers (e.g. the $20 tier), they still have that incentive to burn your tokens and upsell the higher tiers.by nozzlegear
8/23/2026 at 5:37:26 PM
Burning tokens does little to convince people that more tokens are a good use of money - especially if they are choosing between buying more of your tokens and more of your competitor's tokens.by jjk166
8/22/2026 at 10:02:38 PM
the subscriptions all have limits. when you hit the limit you again start paying per token.by blehn
8/22/2026 at 6:19:38 PM
Agreed although the differences between the effort and reasoning is massive. I generally ship 40 hours in three with AI. I could not figure out why my delivery was behind until I started going through the logs. The thinking was extensive, the effort was beyond the original request by a magnitude of 50xby pizzafeelsright
8/22/2026 at 7:14:34 PM
How do you even know that it's consistently 40 hours in 3? What kind of developer ever had that kind of estimation accuracy (unless it's really repetitive) or even focuses on productivity rather than the problem like that? This sounds more like factory work than design or development. I really don't get it.by discreteevent
8/23/2026 at 12:12:11 AM
We are in new territory and not everyone gets it.I am not a developer although I have written production code but that's not what I am talking about.
I know for a fact that I can do a task in a week. It may take 12 or 40 hours but it'll be done in a week. Now with AI? I can get that task done in a day. Anywhere from 3-12 hours.
by pizzafeelsright
8/22/2026 at 8:42:57 PM
> I generally ship 40 hours in three with AI.You ship 3 hours in three with AI. The old number is meaningless now.
by owebmaster
8/23/2026 at 12:07:05 AM
Agreed and management is becoming aware.by pizzafeelsright
8/22/2026 at 7:16:03 PM
Just like me fr frby eli_gottlieb
8/22/2026 at 6:00:02 PM
A.k.a. theft.by chrisjj
8/22/2026 at 7:55:44 PM
nice conspiracy theory, but it doesn't hold. these companies won't last long if people don't get actual work done.by simianwords
8/22/2026 at 8:44:20 PM
There's a big chance these companies won't last longby owebmaster
8/22/2026 at 6:07:14 PM
Thought I was going to be on Claude Code forever.Recently got approved at work for ChatGPT Pro so I could use Codex.
Blown away by the speed. It feels like using Claude Code for the first time again. I don't think Codex is doing anything revolutionary, just better handling of which requests should go to which model, and having faith in some of the "less powerful" models for more than you would think.
It seems the TUI coding experience is very much an open race. This is motivating me to look at other agents / harnesses as well (maybe Gemini, OpenCode, etc).
by rshnotsecure
8/22/2026 at 10:40:41 PM
https://charm.land/crushby andreynering
8/22/2026 at 10:07:42 PM
Codex is faster than Claude, but wait till you use DeepSeek.by a34729t
8/22/2026 at 10:51:02 PM
I initially had unbelievably terrible experiences with Opus 5 and Fable in their higher reasoning levels.I've had WAY better results on medium effort.
IIUC, the consensus seems to be that anything more than medium effort is rarely worth it - and you far more often run into these extreme worst cases than you do with even the lowest effort levels. That definitely coincides with my anecdata.
It's really only worth it if you're hoping to win the lottery asking it to solve an Erdos problem.
by onlyrealcuzzo
8/23/2026 at 2:15:01 AM
I found that I need the big model and the high reasoning effort on tasks where I have a large amount of details of varying levels of importance to keep in mind, all affecting different aspects of the project that might be interrelated to various degrees. Anything less and it would lose track of details. Whereas the very large amount of thinking tokens seem to give the model a chance to "remember" everything it needed to in order to produce good output.For example I'm working on a project now where I need to keep in mind details from 3 separate source repos in different programming languages, along with probably a dozen important business-related documents that either corroborate the stuff in the source repos or add additional important context. It's a lot of details for even a human to manage, and when it comes to actually synthesizing plans and reports across this sprawling information environment, anything less than Opus 5 on Extra+ tends to miss important details and make bad recommendations or draw incorrect conclusions, which then poison subsequent context.
I suspect some kind of RAG-like memory system would greatly facilitate a project like this, but even with such a system I'm not confident that I could get away with an LLM that "thinks" less hard than this. I will say that slogging through the generated documents is kind of miserable and I have to repeatedly fork off side conversations to ask for clarification, but sometimes leads the model to "realize" it's made a mistake in all of its dense babbling, and it's all very hard to interpret. I never had much interest in trying GPT 5.6 until now.
by gwerbin
8/22/2026 at 5:56:14 PM
The business bottom line depends on tokens, shareholders want to see exactly that.by hmokiguess
8/22/2026 at 6:10:43 PM
With how competitive the LLM field is, it would surprise me greatly if any of these players were doing anything other than trying to make the best possible product. I certainly do not believe they are intentionally training the models to use more tokens unnecessarily.by daishi55
8/22/2026 at 6:17:37 PM
Theyre training the system to minimize compute,so most likely theyre dynamically downgrading quants in the first few turns hoping to find the cheapest model to run. The side effect may be excessive token genby cyanydeez
8/22/2026 at 6:28:59 PM
Sure. That’s possible. But that’s not what I was talking about.by daishi55
8/22/2026 at 10:02:37 PM
Tomato, technical tomatoby cyanydeez
8/22/2026 at 8:39:51 PM
There is a theory that the verbosity and comments help getting better results with the current benchmarks. So the models are theoretically getting better but in practice they are getting worse.by owebmaster
8/23/2026 at 2:16:58 AM
I believe they are genuinely getting better for fully autonomous tasks. Anthropic seems to have gone all in on this, at the expense of more typical usage patterns.by gwerbin
8/22/2026 at 6:23:45 PM
Are you implying they can’t do both?by hmokiguess
8/22/2026 at 6:29:24 PM
Yes, those things are mutually exclusive.by daishi55
8/22/2026 at 7:18:16 PM
You're joking, right? That's an incredibly outmoded view of capitalist "competition."by LaGrange
8/22/2026 at 7:48:57 PM
I avoid Opus 5, and reverted to Opus 4.8 over similar issues. Opus 4.8 is still working great for me!by logicallee
8/22/2026 at 6:01:20 PM
and ending with: "One thing I need to tell you:" and bunch of AC, R1 and §by guluarte
8/22/2026 at 5:43:42 PM
Yep, they’re lighting tokens on fire with that thing.by clickety_clack
8/23/2026 at 8:03:48 AM
Even two minutes is crazy long. Does it really takes this long to make small edits with agentic LLMs?by blks
8/23/2026 at 4:33:35 PM
For this specific task there's a good amount of reconciling.by pizzafeelsright
8/22/2026 at 6:11:45 PM
Gotta milk the cows.by brador
8/23/2026 at 6:46:43 AM
They're opitimizing for high token usage so they can charge more money.by phyzix5761