alt.hn

8/24/2026 at 4:30:16 PM

Ox-Alpha Is GLM

https://dejan.ai/blog/ox-alpha/

by jitbit

8/24/2026 at 5:54:28 PM

GLM 5.3 and all previous models don't have a vision encoder and can only accept text. Ox-Alpha can accept video and images, so unless Z-ai added a pretty good vision encoder for this model, I don't think so.

My money is on Moonshot and this being Kimi K3.5. The measured tps and latency is in-line with K3's tps and latency from Moonshot.

MiniMax M3.5 is also possible (but the MiniiMax provider is a lot more performant than the lab behind ox-alpha, so less likely).

by volf_

8/24/2026 at 7:38:42 PM

It would be stranger to me that Kimi switched to GLM's tokenizer than that GLM added multimodal like Kimi and Deepseek both did recently

by nylonstrung

8/24/2026 at 5:55:57 PM

DeepSeek literally just came out with the vision-enabled version of Flash v4 which was purely text based. Why would GLM not be able to do the same thing?

by Almondsetat

8/24/2026 at 6:31:55 PM

It's possible

by volf_

8/24/2026 at 5:58:23 PM

Glm had made vision models in the past. Look up GLM 5v.

The only question now is if it's 5.3v, 5.4/5.5 or a dedicated flash/vision model

by Bolwin

8/24/2026 at 6:31:13 PM

Yeah. It could be. The Z.ai DC latency is still ~1.2s faster than whomever is serving this model.

by volf_