8/12/2026 at 3:38:56 PM
Supposedly this is a Kimi k3 rival. Bit of a chonker, especially since they only released bf16 and fp8. So at launch this will be harder to serve than k3. No QAT on q4 means that someone with deep pockets (nvda?) will have to quant it, with plenty of calibration data. Should bring it ~1.3TB, so around k3 size.License pretty similar to k3 with some caveats. Free to use for internal or <50M$ revenue / year. Limitations above that threshold for serving the model or services targeting coding / productivity agents.
Benchmarks are looking good, trading blows w/ opus4.8 and sol, generally 10-20p under fable. But that's neither here nor there w/ qwen, their benchmark to real world usage correlation has been iffy in the past.
The local model 3.8-27B announced for Friday, same time so ~48 hours from now. That'll be a bit more exciting for a lot more people, since 3.6 was quite good for local inference, and their 3.7-max -> 3.8-max shows a lot of improvement.
by NitpickLawyer
8/12/2026 at 4:08:05 PM
Unsloth already has a guide for their quants: https://unsloth.ai/docs/models/qwen3.8by ZeroCool2u
8/12/2026 at 7:13:06 PM
I had several issues with unsloth gguf, even for models released a few months back like gemma 4, I have 0 confidence in their models, at this stage, I feel several uncensored are more reliable.by Foobar8568
8/12/2026 at 10:38:38 PM
Hey sorry what are the problems that you're experiencing - we're more than happy to help fix them!by danielhanchen
8/13/2026 at 1:05:11 AM
Thank you for what you are doing.by LeBit
8/13/2026 at 2:54:03 AM
Thanks for the support and to the community!by danielhanchen
8/13/2026 at 4:26:36 AM
Counterpoint: I've been using the Unsloth Gemma 4 quants extensively since very soon after release (on ROCm and Apple Silicon, I don't have any Nvidia hardware big enough), pretty much every quantization down to 4 bits (the QAT is the business, indistinguishable from the full-fat version, runs great on a slightly chonky desktop or laptop), and I haven't had any issues. The reason I use Gemma 4 so much often comes down to how reliable it is; when I want to experiment with llama.cpp settings, MTP, n-gram, etc. it's my go-to because I know there isn't anything wrong with the model or the quantizations that could interfere with the experiment.It did take a little while for Unsloth to update the Laguna S 2.1 quants to fix the yarn_attn_factor, and so it was a bit frustrating getting that quantization running right, but almost always, I pick the unsloth quantization if there is one. (Still waiting/hoping for a Ling 3.0 Flash.)
by SwellJoe
8/13/2026 at 7:58:06 AM
We will investigate Ling!by danielhanchen
8/12/2026 at 9:52:22 PM
I've used their gemma 4 quants since when they were not still working in llama.cpp and ik-llama.cpp and I don't remember any problemsThey are the most reliable in my experience, but if you have alternatives you trust I'd love to know
by jokethrowaway
8/13/2026 at 7:49:36 AM
> I feel several uncensored are more reliable.I have had the same experience with gemma 4 on same tasks being refused. But this is when working with cyber offensive tasks and the like. It excels in coding and is very fast on consumer hardware. So I would say use the right tool for the right task.
What uncensored models can you recommend?
by BonerWiener
8/12/2026 at 8:05:45 PM
... because they are often the first to quant it. sometimes the actually model providers will release wrong chat templates or values in the model config which leads to bad quants. how would you know a quant is good if you don't make one? you don't. so they make it first, then they run a lot of tests, KD, perplexity, etc, they publish it. They take feedback from the community, then they update if needed. if you want to try it right now, you grab it else wait for a week or 2.by segmondy
8/12/2026 at 8:16:14 PM
Gemma4 was released a few weeks ago? The problems are still there. Today I started using another "provider" and the problems disappeared. Thanks but no.by Foobar8568
8/12/2026 at 10:39:19 PM
Hey yes - if you could describe what the issues are - we will gladly fix them!by danielhanchen
8/13/2026 at 4:58:05 PM
Mea-culpa, dry-multiplier generated crap and even more so on Gemma4.by Foobar8568
8/14/2026 at 6:12:01 AM
Ok no worries - if there are any future issues - feel free to message / make a HF issue - we'll fix promptly!Also note its best to follow Gemma4's official sampling params since they evaled with it - dry multiplier sometimes works, but it actually screws up reasoning sometimes
by danielhanchen
8/12/2026 at 8:22:12 PM
They are very responsive, and would probably be happy to help you fix your issues.by arcanemachiner
8/12/2026 at 10:39:56 PM
I've 0 issues with gemma4 and I downloaded it early.by segmondy
8/13/2026 at 1:07:00 AM
Same. Used the 12B, 26B A4B and 31B.No issues with llama.cpp.
by LeBit
8/13/2026 at 10:14:26 AM
Without substantiating what issues you have with unsloth gguf files you are just adding noise, no signal.by jacquesm
8/13/2026 at 2:21:05 AM
Huh I have had great luck with unsloth quants so far. What issues are you having?by chlorion
8/12/2026 at 4:33:49 PM
I wonder who is unsloth and where they got time, hardware and knowledge to quantize them?by codedokode
8/12/2026 at 4:50:51 PM
unshloth started as a finetuning library with lots of optimisations so you could finetune on lower end hardware. Kind of OGs of the local community. Started by two brothers Michael and Daniel(?) a math wiz and a community builder/communicator. They've since gotten some VC backing, are active in quantising lots of models on release day (work w/ labs to prepare things), known for their optimised quants (use different bits for different layers). Recently I saw they launched some sort of a desktop app, like lmstudio if you're familiar with it. They're really cool people and known in the local model places.by NitpickLawyer
8/12/2026 at 4:56:37 PM
It's our guy Daniel: https://www.linkedin.com/in/danielhanchenby ZeroCool2u
8/12/2026 at 10:39:45 PM
Hey :)by danielhanchen
8/12/2026 at 4:54:46 PM
They started with offering training methods for quantized models to save memory and added new things over time. They are very active in the local model community and have extensive documentation and tooling to help with running and training models locally.by Eisenstein
8/12/2026 at 6:18:53 PM
Daniel Han is just that good!by arthurcolle
8/12/2026 at 10:39:29 PM
Thanks hahaby danielhanchen
8/12/2026 at 6:52:06 PM
I quantize my models with llama.cpp and it's usually one command. Some of their quants are fine-tuned by architecture but it's only to squeeze out every little performance benefit.by kittikitti
8/12/2026 at 9:47:57 PM
Unsloth imatrix data puts their quants at lower KLD than almost all others.It's true they make architecture-specific changes like keeping certain layers at F16 but it's also more than that.
by suprjami
8/12/2026 at 11:04:24 PM
What hardware would you even be able to run this on?by unleaded
8/13/2026 at 6:47:58 AM
Even the lowliest hardware could run this, but at an unlikely to be useful low speed, e.g. of 3 or 4 tokens per minute (by reading the weights from a couple of 4 TB SSDs for the BF16 model, or from a 4 TB SSD for the FP8 variant).The question about LLMs is never whether they can be run, because that has a trivial answer, they can always be run. The right question is what speeds are achievable for representative hardware configurations.
At launch, it is difficult to estimate the speed. That should be known after someone reports experimental results. Moreover, for many LLMs the speed improved sometimes later after their release, after tweaks in inference backends, like llama.cpp or vLLM.
by adrian_b
8/13/2026 at 9:57:36 AM
tell me how run 35B on my 8GiB VRAM (linux)speed is not problem when You run agents and forget for 2-3 days
by badcafe23423435
8/12/2026 at 9:45:26 PM
The parameter "reasoning_effort" is something new, or am I wrong? Qwen3.8 comes with official support for reasoning_effort, which can be used to adjust reasoning depth and control cost:
- xhigh (default): for complex tasks demanding thorough analysis
- medium: balancing accuracy and speed
- low: efficient reasoning optimizing for speed and cost
In addition, preserve_thinking is enabled by default for all workloads for the best out-of-the-box experience.
Asking because in my case (OCR of scanned historical "National Geographic" magazines) the LLM trying to merge text split into separate columns was running in circles from time to time and needed a lot of prompt tuning when using Qwen 3.0/3.5/3.6 (still needs from time to time).
by zepearl
8/12/2026 at 10:47:21 PM
I'm using Qwen 3.5 for OCR, and reasoning_effort is supported there too. I found that it can be loop-prone (though somewhat less so) even if you set reasoning_effort to low.by philipkglass
8/13/2026 at 11:21:50 AM
From my evaluation[1] neither Kimi K3 nor Qwen 3.8 are as good as GLM 5.2 at coding. I wonder if there's a marketing gap that's got people underestimate it.It uses twice as much tokens to achieve the same but the results are significantly better and because it's so much cheaper it's the most economical choice too.
[1] https://blog.bosun.ai/software-maintenance-with-open-weight-...
by tinco
8/12/2026 at 5:38:08 PM
> 3.8-27B announced for FridayMaybe I’m misreading this or some other post, I thought QWEN was stepping away from releasing these models for local consumption
by alanwreath
8/13/2026 at 10:47:45 AM
https://modelscope.cn/models/Qwen/Qwen3.8-27B/summary Countdown at 28h now.Sadly they seem to not be releasing a sparse 35b A3b or anything inbetween "too large to host for mortals" and "fits into a consumer rtx". Probably not to eat away their profits on their API serving. 120b - 300b is a dead space right now, very few good releases in that size range. (I know there are, but the big labs aren't releasing stuff here)
by underlines
8/12/2026 at 5:44:35 PM
3.8-27B is confirmed for Friday. They didn't release their whole 4B-400B range of models since 3.5. And 3.6 only got 27B and 35B MoE. So yeah, slowing down, but not completely out of the small model game.by NitpickLawyer
8/12/2026 at 5:42:32 PM
They reversed course and now are saying they'll be releasing their Max style models in open weights.by trollbridge
8/12/2026 at 5:43:16 PM
thank you china!by arthurcolle
8/12/2026 at 8:00:45 PM
Rather, thank you competition in China. This is Chinese "overcapacity" (of talent pool) at work.by FooBarWidget
8/12/2026 at 8:20:54 PM
Quite literally the government in China came out and said "It'd be better for us if we did more open models and collaborated with other countries also doing open models" and then Qwen changed their tune. It's literally thanks to China in this case, not competitors/peers in China. Their government is horrible for a lot of stuff, but in this case they do deserve praise for forcing the "right" (according to me) direction.by embedding-shape
8/13/2026 at 1:44:16 AM
> Their government is horrible for a lot of stuff, but in this case they do deserve praise for forcing the "right" (according to me) direction.It's not just according to you.
Without open weights, what happens if you get blacklisted from Anthropic and OpenAI? If AI becomes a standard tool for programming like a compiler, you've effectively been Blackballed from the field of programming. Full Stop. This is "Right to Read" coming home: https://www.gnu.org/philosophy/right-to-read.en.html
In addition, without open source competitors to your core tools, we KNOW what happens. Cadence and Synopsys and a megabuck per engineer per year ... that's what happens.
by bsder
8/13/2026 at 1:21:30 AM
Realpolitik I guess. They want to destabilize the USA, and we want open weights we can run on our own machines. As long as those interests coincide, we are allied.The best outcome for us is the one where they all keep competing and undermining each other until the end of time while providing us all with better models and cheaper hardware to run them with. The US corporations in particular should never be allowed to achieve their "you'll buy intelligence from us on a meter" rent seeking dream.
by matheusmoreira
8/13/2026 at 4:58:35 AM
The US is destabilizing itself plenty on its own. I think this is just a question of common sense, and the party's long standing habit of pushing some competition but not too much competition, to avoid wasted effort.by vintermann
8/13/2026 at 1:26:27 AM
when Mistral finally releases Le Chaton Fat, I will switch to that but until then, I will count my lucky starts that this exists. Cheers to Qwen3.8-2.4T release dayby arthurcolle
8/13/2026 at 4:52:49 AM
Unfortunately they're currently too busy trying to patent tool callingby Zetaphor
8/13/2026 at 6:56:57 AM
Google patented chain of thought too hahaby arthurcolle
8/13/2026 at 1:58:29 AM
Cheers!by matheusmoreira
8/12/2026 at 9:52:10 PM
Our desire for better local models just happens to coincide with China's desire to destroy the western AI company business model by releasing local models. I doubt there's any philanthropy involved.by suprjami
8/13/2026 at 6:49:29 AM
I find it really weird why people keep framing decisions like this in terms of morals and selflessness, e.g., "is/isn't philantropy". This is about relationships and mutual benefit.They literally announced their motivations and world few a few weeks ago at the Shanghai AI conference. They want to ally with the global south. They see AI like the industrial revolution: the global south was left behind for a long time and, as a result has been exploited and has struggled to develop for a long time. They see open AI as a way to level the playing field to prevent such "new historical injustices" (in the sense of the Century of Humiliation and the Opium Wars). Concrete policies to back this rhetoric include technology transfer and training programs for the global south. They frame this latter not as philantropy but as generosity, in the sense that it generates goodwill and what goes around comes around. They believe that helping the global south and cultivating relationships will eventually help China.
Think about it. Your local businesses are not charities either. That doesn't make them bad, nor does it mean you derive no benefit. It still benefits you to cultivate good relationships with them.
by FooBarWidget
8/13/2026 at 8:05:03 AM
China has been using that rhetoric of "helping" the global south ever since it emerged as a super power. In most cases, it has a lot less to do with generosity than with securing natural resources and international influence.by yfontana
8/13/2026 at 11:13:27 AM
> In most cases, it has a lot less to do with generosityNothing happens on a geopolitical scale, from the US, China or anyone else, simply because of generosity.
by embedding-shape
8/14/2026 at 10:19:35 AM
No, but governments will pretend to, and we shouldn't allow them to get away with the pretending.by anshorei
8/13/2026 at 9:00:14 AM
Generosity and securing natural resources and international influence don't have to be mutually exclusive. If you interview Africans, then they say that while cooperation with China has problems, they sure are glad they at least have more choices now, and on the whole China's existence is beneficial. The alternative — only western choices and no China — sure hasn't served them well in the past half century.No matter what you believe is their "true" intentions, offering 5000 training and tech transfer positions to the global south is a very concrete and unambiguous move. As are forgiving African loans and unilaterally offering zero trade tariffs.
by FooBarWidget
8/13/2026 at 1:01:37 PM
If what China has been doing is considered generous, then so should the tens of billions in foreign aid that the West has sent to the South over the years.I'm not saying that China's investment in the South hasn't had positive effects. But "generosity" is rarely a relevant lens when analyzing international relations.
by yfontana
8/13/2026 at 1:24:46 PM
> then so should the tens of billions in foreign aid that the West has sent to the South over the years.Well, yes? Why does it have to be either-or?
> But "generosity" is rarely a relevant lens when analyzing international relations.
Automatically assuming nefarious intentions behind all moves is also rarely a relevant lens.
And as I said, and I'm not sure why you keep ignoring it, but I define "generosity" in the sense of mutual benefit. Being nice to your neighbors and helping them, benefits you due to generated goodwill. I'm explicitly not defining generosity in the sense of selfless philanthropy where you get nothing back. The idea that doing good things for others eventually results in good things coming your way, and thus that one should do good things for others even it's selfishly motivated (and also that there's nothing wrong with this), is not a crazy idea.
In a lot of cultures (Chinese included), gifts are not simply gifts. There is the social expectation that the gift is reciprocated. Western cynicists may call this "manipulation" or "influence". The Chinese see this as the start of a relationship of a cycle of mutual gift giving.
by FooBarWidget
8/13/2026 at 12:29:27 AM
Which coincidentally helps average people far more than the already rich investors in a few western mega corps. If it weren’t for these big Chinese model releases the western companies wouldn’t release anything at all. The field would be advancing at a snail’s pace.by horacemorace
8/13/2026 at 9:20:09 AM
Its not just that.China also puts pressure on rich chinese flaunting their riches.
They have a common prosperity initiative.
by Manfrednotfunny
8/13/2026 at 6:40:22 AM
And what do you think happens after the govt said that? Going to companies and force them to open models? That's only how westerners' misconception of China works. If you study Chinese EV industrial policy history you'll know they're not based on coercion of private parties but on incentives (and that, ironically, private EV companies succeeded despite incentives, not because).The real policy mechanisms around open AI models are also incentives. Various cities have programs to pay companies for releasing open models. They subsidize compute through vouchers. They reward universities and students for open source collaboration.
This isn't some black box. The policies are written down, anybody can read their AI+ policy papers.
Alibaba went back to releasing open models way before the Xi speech from a few weeks ago. The cause is pressure from researchers, who believe in openness, as well as the competition who keeps releasing open models. This is Chinese "involution" at work. And the subsidies also help, of course.
by FooBarWidget
8/12/2026 at 9:54:40 PM
The obvious goal is to destabilize the western economy and prove that US tech is a worthless bubble - but I agree, OSS AI is great for everybody and what OpenAI was supposed to beby jokethrowaway
8/12/2026 at 11:48:56 PM
There’s an alternate universe in which OpenAI stays open, licenses according to revenue, Chinese models don’t gain traction in the US because domestic models take all the capacity…whatever, $1T IPO beats the right answer ever timeby DrBenCarson
8/13/2026 at 2:50:10 AM
> Chinese models don’t gain traction in the US because domestic models take all the capacity…whatever, $1T IPO beats the right answer ever timeThe trouble here is how more infrastructure helps OpenAI and Anthropic continue billing at 10/100x Chinese model rates.
Either their models have to be better (to justify the higher prices and margin) or their inference has to be lower cost (which isn't going to happen until they move away from Nvidia).
by ethbr1
8/13/2026 at 7:00:07 AM
Black market operators can resell stolen account tokens at a lower price than authentic premier tokens from frontier labs and can host their own infra too. I'm not super convinced frontier model serving without downstream model development on a vertical specific software / knowledge worker "factory" model can workby arthurcolle
8/13/2026 at 3:42:26 AM
that "collaborate" means US GPUs and training set, then deployed in censored data centers for profitby ngl999
8/19/2026 at 6:24:15 AM
sorry I can't hear you over all the tokens getting sprayed all over my workstation, bubble party styleby arthurcolle
8/12/2026 at 8:13:45 PM
[dead]by PerkFuel
8/12/2026 at 3:53:31 PM
Now that they have reached the frontier in raw performance, I would like to see Chinese models improve their reasoning efficiency.by esafak
8/12/2026 at 10:01:30 PM
I bet they could do it faster if they weren’t blocked from buying GPUsby JSR_FDED
8/12/2026 at 5:29:57 PM
For all the talk about over reasoning, K3 on low thinking has been rather niceby verdverm
8/12/2026 at 6:27:25 PM
It looks like a work horse! Is it yours?by esafak
8/12/2026 at 3:57:44 PM
[flagged]by jingpostmedia
8/12/2026 at 3:53:17 PM
Llama.cpp can quantize without special training, but I'm not sure if any special model architecture support is needed to read it in the first place. If it can be converted to gguf at all and you know what tensors to target, it can get the full ternary bonsai treatment today.by MrDrMcCoy
8/12/2026 at 4:10:28 PM
Sure, but that's for "personal" serving. I meant for 3rd party providers. Usually we get a good indication on what it costs to host this, as the prices settle on open router. That's why I said it's tougher to serve than kimi k3 on launch. As a provider you'd do fp8 if the model creator didn't do QAT on q4, or until someone does a good calibrated nvfp4. And that's usually nvda :)by NitpickLawyer
8/12/2026 at 4:19:51 PM
That makes sense, but your specific phrasing precluded the possibility of non-QAT quantization.by MrDrMcCoy
8/12/2026 at 4:24:59 PM
Should have worded that better, my bad.by NitpickLawyer
8/12/2026 at 3:59:07 PM
QAT is an optimizing quantization algorithm, not naive quant.by binary132
8/12/2026 at 11:56:42 PM
Isn’t QAT a training approach (roughly, simulating quantization in the forward pass during training so that quantization of the level targeted in training has close-to-optimal behavior), not a quantization algorithm? Hence, the name?by dragonwriter
8/13/2026 at 6:42:07 PM
sounds like a repeatable method of optimizing quantization to me, friendby binary132
8/12/2026 at 4:05:49 PM
Right, but the way they phrased it suggested that without QAT it could not be quanted at all.by MrDrMcCoy
8/13/2026 at 9:16:47 AM
Show me any open source model quantisite, distile or reduce size from nvidia ;)by badcafe23423435
8/12/2026 at 5:28:48 PM
quanting is actually cheap and you can compress a model that does not fit on a GPU. You can process layer by layer, this is what the sequential processor in llm-compressor does.by verdverm