Kimi K3, Qwen 3.8, and Anthropic's (potential) Unravelling
Kimi K3, Qwen 3.8, and Anthropic's (potential) Unravelling
Kimi K3 and Qwen 3.8 prove open models can reach the frontier. Why the economics of frontier labs favor infrastructure owners — and why Anthropic's model-only position risks unravelling.Wojciech Gryc (Emerging Trajectories)

warmaster
in reply to ☆ Yσɠƚԋσʂ ☆ • • •Dumb AF acronym.
Eager Eagle
in reply to warmaster • • •common English phrase
Contributors to Wikimedia projects (Wikimedia Foundation, Inc.)iturnedintoanewt
in reply to ☆ Yσɠƚԋσʂ ☆ • • •☆ Yσɠƚԋσʂ ☆
in reply to iturnedintoanewt • • •iturnedintoanewt
in reply to ☆ Yσɠƚԋσʂ ☆ • • •relic4322
in reply to ☆ Yσɠƚԋσʂ ☆ • • •There was never a question of if open models can match a frontier. Open just means you released the weights. The question is how they trained it. Most of the models coming out of China are distilled from larger commercial frontier models, and those are not the same thing. More situationally brittle.
Can be really good in narrow lanes, but not the same category at all. And that doesn't show up in benchmarks.
☆ Yσɠƚԋσʂ ☆
in reply to relic4322 • • •People really need to stop parroting this line uncritically. The process takes time because even when you're distilling answers, you still need to actually do reinforcement training on the model. And given that Fable and GPT 5.6 just came out there simply hasn't been much time to do that. However, models like Kimi also do better than Fable or GPT on a lot of tasks, which means it's not just distillation but also difference in architecture. You can watch to see how Kimi was actually trained and why it performs well.
It's also absolutely hilarious that people think only Chinese companies use distillation, as if Anthropic or OpenAI are above that or something. Not to mention that they basically ignored copyrights on all the data the siphoned and are now crying that people aren't respecting their terms of use.
The reality is that
... Show more...People really need to stop parroting this line uncritically. The process takes time because even when you're distilling answers, you still need to actually do reinforcement training on the model. And given that Fable and GPT 5.6 just came out there simply hasn't been much time to do that. However, models like Kimi also do better than Fable or GPT on a lot of tasks, which means it's not just distillation but also difference in architecture. You can watch to see how Kimi was actually trained and why it performs well.
It's also absolutely hilarious that people think only Chinese companies use distillation, as if Anthropic or OpenAI are above that or something. Not to mention that they basically ignored copyrights on all the data the siphoned and are now crying that people aren't respecting their terms of use.
The reality is that China tops the world in artificial intelligence publications today. Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer. Anybody who thinks China is simply distilling glorius American models is not engaging with reality.
This whole narrative has just been a massive cope.
- YouTube
www.youtube.commarcie (she/her)
in reply to ☆ Yσɠƚԋσʂ ☆ • • •☆ Yσɠƚԋσʂ ☆
in reply to marcie (she/her) • • •relic4322
in reply to ☆ Yσɠƚԋσʂ ☆ • • •I didn't say that. I described the difference and stated a fact about distilled models vs frontier coming out of China. When you are distilling from a source without permission you get degraded results.
I also didn't say that distillation is bad, I said they have narrower use cases. Nor did I say american companies are great.
But, please continue throwing your assumed biases out there.
☆ Yσɠƚԋσʂ ☆
in reply to relic4322 • • •