Global downloads of China's open-source AI models exceed 10 billion
Global downloads of China's open-source AI models exceed 10 billion: official
China’s innovation capability continues to strengthen since the beginning of 2026, with the penetration rate of AI technology across the country’s major large and medium-sized enterprises exceeding 30 percent, and the cumulative global downloads of C…www.globaltimes.cn

i_am_not_a_robot
in reply to ☆ Yσɠƚԋσʂ ☆ • • •☆ Yσɠƚԋσʂ ☆
in reply to i_am_not_a_robot • • •ms.lane
in reply to ☆ Yσɠƚԋσʂ ☆ • • •☆ Yσɠƚԋσʂ ☆
in reply to ms.lane • • •brucethemoose
in reply to i_am_not_a_robot • • •You can run MiMo 2.5 in 128GB, as 3 bit quant. The model itself is less that 128GB as an IQ3_KT, and it’s KLD (measured quantization loss) is very reasonable.
I get about 9 tokens/sec on 128GB with a 7800 CPU and one RTX 3090. Intend to try dflash to speed it up this week. I can run up to 90k context, without too much kv quantization, depending on how I configure my PC.
And it’s a really good model, even quantized.
chicken
in reply to ☆ Yσɠƚԋσʂ ☆ • • •userentity
in reply to ☆ Yσɠƚԋσʂ ☆ • • •Qwen 3.6 27B MTP (Q4) is really great. ~24GB VRAM usage with one slot @ 262k . ~30GB with 3 slots at 128k. Also it does not struggle like Gemma if you use e.g. a Q4 KV-Cache. And it runs at 400-800 ppt/s and 20-60 itp/s on a V100.
But it has a competitor since a short while that is Laguna XS 2.1, which is really good for a A3B MoE. I'd never have thought a 30B MoE could be on par with a 27B dense, but it seemingly is.