Skip to main content

in reply to mynameisbob

Sir, this is a ~~Wendy's~~ technology forum.
This entry was edited (6 days ago)
in reply to P03 Locke

Indeed, and as Cory Doctorow says: “The most important thing about a gadget isn’t what it does; it’s who it does it for and who it does it to.”

::: spoiler Tap for spoiler
Fuck AI
:::

This entry was edited (5 days ago)
in reply to ☆ Yσɠƚԋσʂ ☆

And the latest Kimi is better than Fable 5. Another plane has indeed hit the Anthropic towers.
in reply to Communist

It does not seem to be better


It does.

except in terms of cost


Except cost. Except open source. Except the entire fucking global economic trillion dollar US AI model.

Those are very big exceptions.

(Also, I'm mostly stealing from Yσɠƚԋσʂ's post.)


Just a day ago, a senior Anthropic executive claimed the U.S. still held a 6 to 9 month lead in frontier AI models, while calling Chinese model distillation adversarial. One day later, Moonshot’s Kimi K3 beat Claude Fable 5 on Frontend Code Arena. And the funniest part is that it is an open weight model.

So the whole 6 to9 month lead lasted about 24 hours. 🤣

hai.stanford.edu/news/inside-t…


in reply to P03 Locke

That's one benchmark that they focused on, but having double checked, you're right, I was thinking of this one where it lags behind gpt 5.6 deepswe.datacurve.ai/
in reply to P03 Locke

I'm just waiting for the closed-source AI industry to have their Black Swan moment.

Something's going to come out of left field from the open AI community and all that investment in proprietary models and mega-compute is going to be rendered useless.

This entry was edited (5 days ago)
in reply to Dave.

You still need the mega-compute. Even for inference, 1.5 terabytes of RAM in modern servers isn't cheap.
in reply to eleitl

You can run Claude Sonnet equivalent quantised models locally on much less RAM.

Local LLMs are close to being viable. We're almost at the point where they fit a Macbook.

in reply to Grapho

Not really. I’m glad that China is copying AI models and providing them for cheaper usage or open sourcing the training.

I don’t believe any US company should monopolize the entirety of human knowledge.

in reply to nosuchanon

I gotta know what you mean by “copying” ai models. Same training data? Amount of training? How would you go about copying a closed-source model. I think china is just…. Making good models.
in reply to ResistingArrest

latimes.com/business/story/202…

Distillation is a technique where an older “teacher” AI model is used to train a newer, “student,” model that replicates the capabilities of the earlier system — often at a much lower cost than producing an original model from scratch. Some forms of distillation are widely accepted and even encouraged by AI labs, such as when companies create smaller, more efficient versions of their own models, or allow outside developers to use distillation to build non-competitive technologies.


Basically one AI asks questions and records the answers to recreate training data without massive compute for training data.

in reply to nosuchanon

I see and agree with you then. I’m glad it gets us better cheap models but more traditional innovation would definitely be preferred.
in reply to nosuchanon

You can't distill models that fast lmfao. People just keep parroting this without having any clue how any of this works.
This entry was edited (5 days ago)
in reply to ☆ Yσɠƚԋσʂ ☆

It’s not that complicated. Recreating training data via distillation is basically asking structured questions and recording the responses and reformatting that to use as cleaned “good” training data. Much less energy and compute intensive than creating the training data on your own.

I think I remember reading somewhere how Chinese research’s do this by basically using bots and spreading out the distillation to many source queries.

This entry was edited (5 days ago)
in reply to nosuchanon

The process takes time because even when you're distilling answers, you still need to actually do reinforcement training on the model. And given that Fable and GPT 5.6 just came out there simply hasn't been much time to do that. However, models like Kimi also do better than Fable or GPT on a lot of tasks, which means it's not just distillation but also difference in architecture. You can watch to see how Kimi was actually trained and why it performs well.

It's also absolutely hilarious that people think only Chinese companies use distillation, as if Anthropic or OpenAI are above that or something. Not to mention that they basically ignored copyrights on all the data the siphoned and are now crying that people aren't respecting their terms of use.

in reply to ☆ Yσɠƚԋσʂ ☆

Yeah I know they’re not the only companies doing distillation, it’s just currently in the news and on peoples minds.