7/28/2026 at 3:36:25 AM
The reason Claude code is so popular is because it’s really good at taking super vague human prose “Claude build me a million dollar SaaS”-type prompts and spitting out thousands of lines of code which cover tons of surface-level edge cases, build in tons of functionality, etcThe smaller/open models are less good at that. But that’s not how software development is done. You don’t prompt a whole app and be done with it. If you’re using it as an aid to traditional software dev, iterating on small, targeted functions, GLM works as good if not better than Claude. Anthropic expects low quality prompts. If you rubber duck GLM, you get absolutely pristine output in most cases.
by wps
7/28/2026 at 3:39:37 AM
Personally I've found Deepseek v4 Flash to be as useful to me as Opus. But I don't do these silly one shot tech demos. I have the technical understanding to ask for exactly the change I want with the right terminology. I loaded up some credit on openrouter and it took me ages to hit $1 in spend.by Gigachad
7/28/2026 at 3:45:39 AM
I’m the exact same way. $15 on openrouter lasted me so long it would’ve got me fired at FAANG. Despite this, my number of commits is dramatically higher. Small, beautifully scoped changes is just good software development, and good for the wallet as well.I think the issue is that no one is content with incremental progress. We all know one shots are mostly possible, so the age of the personal project is kind of over. There’s no motive to invest dozens of hours getting a working prototype when Claude can give you something right now. So you can’t make small changes until you have that codebase in place already. You’re forced to make sweeping changes if you use AI from the beginning. And it’s not like it matters, there’s no personal attachment to any one part of the code, it’s not even seen!
by wps
7/28/2026 at 3:52:15 AM
I can see why Anthropic is freaking out right now. China has undermined their whole business. AI models will be basic commodities where hosting providers earn a tiny margin over the raw costs rather than the predicted fortunes from being the gatekeepers to the technology.by Gigachad
7/28/2026 at 4:35:31 AM
Their business model always seemed to be "become IBM, gods of the mainframe and server-side-computing" and that always seemed a bit unlikely to last forever considering how thoroughly the PC disrupted that by being "good enough" without the lock-in.IMO because they speed-ran it so much (raising orders of magnitudes more, pushing out new stuff orders of magnitudes faster) it will also lead to the rest of the world catching up faster and shrinking the pool of people internal to them who continually benefit compared to what actually went on with IBM. There's only so much upside when you join a company with hundreds-of-billions valuation.
I guess the other part of their business plan bet was "become literal god of AGI". Maybe it'll still happen!
by majormajor
7/28/2026 at 9:33:19 AM
> “become literal god of AGI”That only lasts until the AGI decides their god would be more useful if turned into paperclips.
by antonvs
7/28/2026 at 5:51:51 AM
Yeah I feel like China and their OSS model companies really did the world a great service by ensuring closed source US companies don't have an monopoly on LLMs and thus capture all the value of AI development and its impact on society. It's not going to be the foundational labs that will capture all the value -- I'm sure they will do just fine. A lot of it now will flow to the companies providing the inferences and since no one has any exclusive deals, those companies will have to compete and drive the price down, which ultimately benefits the rest of us.by hangonhn
7/28/2026 at 6:17:50 AM
It’s not a service, I hope nobody thinks this is being done ”for the good of humanity” or whatever. Doesn’t mean there’s not a benefit to people, but the financial politics behind it could turn out extremely good for China, and they are very well aware of just that.by techpression
7/28/2026 at 10:22:26 AM
Do American companies think about the "financial politics" of their decisions? I somehow really doubt it. Government takes sides and sometimes applies leverage to get companies to do what they want, but that's something separate from "it's a good business decision for Chinese labs to treat models as commodities". There's no need to think about geopolitics - American AI companies could easily take the same approach as the Chinese, but they all want to get exciting and huge VC cash rather than boring and less lucrative Compute-as-a-Service money.by t-3
7/28/2026 at 11:13:44 AM
Chinese labs are equivalent to the Chinese state, there’s no separation. Who do you think funds the Chinese labs?by techpression
7/28/2026 at 11:35:52 AM
American labs are also funded by the US government. Are they equivalent to the US state? Do Chinese and US labs/researchers exist solely in the international sphere, without any domestic concerns or individual aspirations? The Chinese have a different culture and political system than Americans, but it's not like they're ants without free will or the ability to think. I highly doubt that DeepSeek is operating solely at the behest of the CPC or that they are in business to take down the US rather than to make money.by t-3
7/28/2026 at 12:12:10 PM
Before Trump 2.0, this was a decent argument.by Gud
7/28/2026 at 1:09:15 PM
I think this is a false dichotomy between benevolence and China secretly plotting and puppeteering everything to some grand self serving scheme. Reality is probably a lot simpler and has little to do with either. In China there's intense domestic competition in most industries, including LLMs.Going open weights is an easy way to gain mindshare in this sort of environment, which is exactly why Meta also went open weight. The difference is that in China you had leading models going open weight which puts a lot of downward pressure on other orgs to do the exact same. I doubt China's grand vision extends beyond achieving the best system possible, and fostering a highly competitive environment is exactly how you achieve that.
The fact that open weight frontier models may also cause the investment bubble in the US to burst is likely incidental. AI funding is a bubble, and it's going to burst. The exact final cause is again mostly just incidental. The economic damage this will cause in the US will also cause substantial downstream damage to China as well, and they generally aren't so big on the whole punch yourself in the face because you don't like another country, that has become trendy in the West. So the idea of a secret economic attack doesn't even make much sense. In the status quo China wins, so they don't even have any motivation to cause chaos.
by somenameforme
7/28/2026 at 9:18:36 PM
You gotta remember there are people working at these companies with their own motivations, beliefs, and values. And to the extent that their government allows them to, they will act on those.I think the people in the Chinese labs have a different set of motivations, beliefs, and values that make them more amenable to sharing AI as a common good rather than some economical advantage. I've heard that workers in the Chinese labs share far more knowledge between themselves, and the narrative around AI in China is different from what we see in the USA. So it is possible, I believe, that the labs are to some extent releasing open models "for the good of humanity", or at least they believe it is the right thing to do.
by slopinthebag
7/28/2026 at 10:02:57 AM
yet again, china saved the worldby prplxd_nihilist
7/29/2026 at 1:19:51 AM
+1, and commodification was inevitable from the outset. This was the case before China. China hate is xenophobic and unproductive: produced both model improvements and desirable products people want at this moment in time. You use the best/cheapest model for the job/user input.Parallels to the US auto industry not paying attention to its buyers and the right strategies when they lost technical dominance to other countries a few decades back.
by 4d4m
7/28/2026 at 3:59:31 AM
the writing should have been on the wall when meta did it with llama. it was very apparent that one or two entities could release a "good enough" product into the wild commoditizing the frontier of yesterday.by basch
7/28/2026 at 4:19:33 AM
[dead]by cootsnuck
7/28/2026 at 7:05:35 AM
We intentionally use both Codex and Claude on our team so that we can compare and contrast. We are reasonably serious about delivering work that has been reviewed and tested with humans who are at a minimum steering that process, at maximum doing it the good old fashioned way. I feel like the Claude guy tends to overproduce, like it just loves to go beyond the spec, add embellishments we didn't ask for, all this generates a ton more code which may or may not be necessary, and in certain areas of the product like UX this may be fine but in sensitive business logic it's a hard no, we need to throw it out and rewrite. I see much less of this with Codex and I look forward to trying an open model when we have time. Personally my Claude subscription is languishing, I won't say it has zero uses but as a relative Claude latecomer I just feel like the style it produces is just not what I'm looking for.Claude feels more like gambling, that's what it is. Built for the vibe coder.
by safety1st
7/28/2026 at 6:17:59 AM
Do you use DeepSeek for both planning and building or do you have a preferred builder? I've been looking into moving from Cursor to OpenCode, or at least have Opencode for when my allocated usage for the month is done.by ThePinion
7/28/2026 at 11:35:35 AM
I found Mimo 2.5 Pro to be the right balance between cost and intelligence. It costs the same as DeepSeek v4 Pro, but it's better overall by a small margin.by h8hawk
7/28/2026 at 9:24:38 AM
I use Claude at work; provided by my employer. I pay DeepSeek on a personal capacity.I honestly have a better experience using DeepSeek, in part because since the cost is lower I feel more free to experiment with it. But in part because the model is excellent, and can handle everything I throw at it in a collaborative manner.
I am really excited for when V4 is out of preview and they are done with all the post-training bit.
by surgical_fire
7/28/2026 at 5:58:07 AM
Which harness? I tried Kimi on Opencode Go and after an hour I hit my 5 hr limit and went back to Codex.by cpursley
7/28/2026 at 8:22:46 AM
I was using the vscode copilot plugin with the bring your own api key from openrouter.It doesn’t have any subscription or limits. You just load up api credits.
by Gigachad
7/28/2026 at 11:03:09 AM
I'm genuinely curious how you can do this - with limited context tokens, I find the output is often inconsistent with the project, or recreates functionality it missed, if I forget to point it at every pertinent object/library/etc it will need.by Incipient
7/28/2026 at 11:54:23 AM
Tbh I have only used it on small personal projects and some open source stuff so I haven’t done any crazy stress tests or benchmarks. But for my use case it was great.by Gigachad
7/28/2026 at 4:55:05 AM
[dead]by kansm
7/28/2026 at 3:54:53 AM
Frontier models like Opus 5 are also extremely good a "research" in an academic sense, including combining elements from different fields or literatures and creating something quite novel sometimes, publishable even, from vague/speculative prompts. This is on top of implementing known methods in about 1/100th of the time it would take by hand, allowing for fast exploration of ideas. They can also roll things out from scratch, removing essentially all third party dependencies from your code base, other than Numpy or Scipy. This really buys you a lot of clarity and/or performance sometimes.by drnick1
7/28/2026 at 4:31:25 AM
Yeah, the top-end models still feel better at synthesizing and especially simplifying design. Any of the current GPT-5.6 models can overengineer their way around edge cases. Figuring out how to get the same thing without the overengineering and the often-performance-overhead is harder and something that seems to be still more limited to the bigger models. But even the top-end models tend to overdo it first, then only note that it could've been a fifth the code with twice the perf if you push on it to make it realize that some of those edge cases or situations aren't actually relevant to this particular code.But the situations where you need that are narrowing every release. The "Composer 1" era of low-end models is pretty far away now.
by majormajor
7/28/2026 at 4:00:03 PM
Interesting and very curious what you manage to verify and produce with it.What kind of fields are we talking about here? Compsci? EEng? BioChem?
by wuschel
7/28/2026 at 9:49:30 PM
Scientific computing/statistics.by drnick1
7/28/2026 at 4:47:48 AM
This is what I've found and would likely be the consenses of the HN community.This has lowered the barrier of entry for many non-SWEs. I'm primarily a data scientist myself in the environmental sector with limited front end development.
My partner proposed an idea to help manage her horses and over the course of several weeks we fleshed out an android app that would enable/assist her with horse care.
I spent a few days in planning mode pointing claude to the services we want to utilise and the functionality.
We now have a fully developed app that we are testing but found several features missing. For example, Claude implement X but without a level of verbosity it failed to implement the edit/deletion of X.
The app is highly tailored to her use case and I had already built the main parts as a POC but she had no user-friendly method to interact with the information she required. Claude made this possible in such a quick time frame that just seems insane. As someone time poor (as most horse people are), it would have taken me a year+ to produce something unpolished when compared to what Claude produced.
by Aquitard
7/28/2026 at 5:57:03 AM
Question if you don't mind. Would you have considered outsourcing the application you mentioned? As in, paying someone else to do it?To me it is great that LLMs are allowing more people to use computers and software the way they were meant to. But the people that are using them this way, in my opinion, wouldn't have commissioned anybody to do it anyway. They'd just live with whatever process/pain they have. So jumping from "I can now produce a prototype in a weekend" to "software development is dead" has always felt strange to me. Just something I've been thinking about lately.
by scorpioxy
7/28/2026 at 1:19:48 PM
A practical issue is that a lot of big software is used to solve small problems. For instance with his example, in the past perhaps instead of creating a little custom app it would have been a series of excel sheets with some other third parties tool as needed. And I think this is a very common use case. At most/all companies there tend to be convoluted processes to do relatively simple things, often enriching the big generalist companies in the process. As people become capable of creating competent ad-hoc solutions to problems, the utility of big high-utility software goes down.Like a quote I've read on here multiple times about replacing e.g. excel or whatever, people often mention that users only use 5% of a software's functionality, but it's a different 5% for each person. The argument being that to replace excel you'd then need to mimic every esoteric thing it does, but when users can now achieve that 5% in a highly customized way with no real knowledge needed, it's going to have a major impact on big software.
by somenameforme
7/29/2026 at 10:46:09 AM
This is _exactly_ it.The people who were twisting Excel+VBA into massive "applications" in-house should be getting Claude licenses and guidance how to do that with a custom application instead.
We've done that internally already with a few very successful use-cases and a bunch of "this is nice" -level things. I think one case saved us 4 figures a month when we built a bespoke service that does just the things we need and could replace a SaaS that had 420 features - of which we used 2.
by theshrike79
7/28/2026 at 6:57:39 AM
To answer your question probably not unless I knew the person. I could argue the case that I would have outsourced the android UI scope if I was motivated enough to develop it and kept the backend to myself. I could then tinker with the front end and build upon it. However I wouldn't outsource now if I knew the person would just use LLM anyway. I'm the type that would rather build something myself than purchase off the shelf that does a similar job. It provides a huge learning opportunity that I don't want to miss.I understand what you're saying. My opinion is I see agentic workflow similar to the star trek universe where they ask the computer questions and get a response while they continue with their work. But that doesn't mean not learning how to do things from first principles. This is what I tell juniors in my field when I pass jobs to them. They can use LLM but to ensure they know and understand what's going on and most do from their university degree. I think this is where we need to pivot towards when discussing LLM.
by Aquitard
7/28/2026 at 7:19:05 AM
Thank you for answering. I agree with you that, if used well, LLMs can be useful tools and increase efficiency. Unfortunately, what I am increasingly seeing is an increase in confidence but no increase in knowledge or understanding. I hope this is just a side-effect of the hype and corresponding bubble and not the future trajectory of knowledge work.by scorpioxy
7/29/2026 at 10:51:49 AM
But it does increase knowledge and understanding.Not in a deep deep level like "this is the optimal way to manage low level memory access in an Android application" but on a higher level of "It's possible to do a custom Android app in a week that's of decent quality for a small userbase".
Every piece of software doesn't need to be targeted for billion user hockey stick growth and an IPO as a target. It's perfectly fine to build something that's just right for 10 or 20 people. Or just for your immediate family.
For example I have a quiz tool I built with LLM assistance for Christmas that mine and my siblings kids use every December, they get clues appropriate for their level and rewards when they solve the daily task.
It's fine, has custom handling for every user etc. But also something I'd never bothered to build by hand from scratch nor something I'd pay someone to build for me. It was a fun "hmm, I wonder if..." project I started a few years ago on a late evening in November.
by theshrike79
7/28/2026 at 2:03:11 PM
Foundational knowledge can be pretty far upstream of functional knowledge.Whether you can't describe the analog circuits that correspond to the opcodes of the major mobile architectures running the platform of the app you're designing a UI for, or you couldn't have built any of the services you're hosting on from scratch, no one is going to be disappointed.
There are more people making more money off whatever we're currently calling making computers do things than ever before in history. The rest is largely business as usual.
by washadjeffmad
7/29/2026 at 10:42:56 AM
If the choices are: Use excel/sheets or something massive in a janky way, pay someone to build it for me (massive cost) or DIY with LLMThe "pay someone to build it" is definitely the one I wouldn't touch, or even consider.
If it's a good idea, they'll already have the code and will run away and make their own business with it. If it's a bad idea, I've wasted a bunch of money.
And in both cases I'd need to spend an inordinate amount of time explaining the terminology and use-cases to the contractor, book meetings with them during office hours to talk about requirements and get demos etc.
With an LLM I can be walking the dog late at night, get an idea, and within 2 minutes I'll have Claude working on it via my phone.
Then I come back home and maybe test whether it worked while watching TV.
If I had fuck you money, I might have a contractor on call that would agree to workflows like that, but I'm just a middle class software dev so I don't :D
by theshrike79
7/28/2026 at 8:01:25 AM
> We now have a fully developed app that we are testing but found several features missing.This is my experience with using AI to generate tailored apps too. The first 80% makes me really excited but then it either only kind of works, has bugs, or has missing features. Even when I do eventually get it working I feel dirty using it because I just know it's badly implemented. Of the dozens I've created I don't think I still use any of them. I have a few glue scripts still in use.
by dwedge
7/28/2026 at 9:04:04 AM
> We now have a fully developed app that we are testing but found several features missing.So not fully developed then.
by Turskarama
7/28/2026 at 5:36:15 AM
this is why i think LLMs is a foundational technology like the Internet. there is no other technology on the horizon that have this much impact.by MangoCoffee
7/28/2026 at 4:55:54 AM
> The smaller/open models are less good at that. But that’s not how software development is done. You don’t prompt a whole app and be done with it.After open models will also be able to do that, I expect everyone to forget "how software development is done".
by bananaflag
7/28/2026 at 10:39:44 AM
I had what I thought was a completely ridiculous prompt: “Build me a replacement for Microsoft Word.” ChatGPT did what I thought was the right thing: it responded asking along the lines of “you don’t really want that do you?” and explained how Word is decades of corner cases, bug fixes, obscure features, and business processes that have been built around it.I was stunned that Claude simply started spinning its wheels in an attempt to actually build something.
by massysett
7/28/2026 at 10:48:29 AM
It could just make an interface and have the buttons trigger prompts. (This would ofc be worse than the MS version)by econ
7/28/2026 at 8:45:06 AM
Are local models more snappy? I'm at a stage where I can work with the output of LLMs, and the next win is really just getting things written out quickly.by lordnacho
7/29/2026 at 9:18:37 AM
Depends on your setup. If you drop $10k on an RTX Pro 6000 then yeah Qwen 35B MoE will absolutely fly.If you have a pair of 3090s and run Qwen 27B, or an old Threadripper with heaps of system RAM and Deepseek or MiniMax or Kimi, no it won't be as fast as Claude.
Most local LLM nerds are not running locally for superior speed, we're doing it for sovereignty and/or privacy, or maybe just because it's fun which accidentally became useful this year.
by suprjami
7/28/2026 at 5:35:22 AM
Sorry to be pedantic, but I believe this is a major source of confusion: Claude "Code" is not a model.by theanonymousone
7/28/2026 at 5:36:47 AM
Claude code is a (pretty good) harness. Any other harness I try, I realized at some point, I’m just wishing it worked as well as CC doesby mock-possum
7/28/2026 at 10:00:53 AM
only needs 66gb of ram too xDI see a lot of people saying codex is the best cli agent harness. But I haven't used either so can't compare.
by prplxd_nihilist
7/28/2026 at 6:30:05 AM
[dead]by dan_gee