LongCat-2.0: China Breaks Free From Nvidia to Train a 1.6T Model

LongCat-2.0: China Breaks Free From Nvidia to Train a 1.6T Model

More

Summary

Meituan — better known outside China as a food-delivery giant — has quietly released LongCat 2.0, a 1.6-trillion-parameter mixture-of-experts language model that activates roughly 48 billion parameters per token. What makes the release particularly notable is that the entire pre-training run, spanning more than 35 trillion tokens, was completed on domestic Chinese AI supercomputers with no Nvidia GPUs, no rollbacks, and no catastrophic loss spikes — a meaningful signal that frontier-scale training is no longer exclusively dependent on US chip supply chains.

The video walks through LongCat 2.0’s architecture, which the team calls LongCat Sparse Attention — an evolution of DeepSeek’s sparse attention approach that attacks the indexer bottleneck from three angles simultaneously: more predictable memory access patterns, shared indexing passes across neighboring layers, and a coarse-then-fine scoring strategy. On benchmarks, LongCat 2.0 trades blows with Claude Opus 4.6 and edges out Gemini, though it falls short of the very latest closed frontier models like GPT-5.5 and Opus 4.8 on the hardest agentic tasks.

Host Fahd Mirza demonstrates the model live on Meituan’s hosted interface, prompting it to generate a complex self-contained HTML file simulating a global railway hub — a task that tests long-context instruction following and code generation. A multilingual translation test across 80-plus languages rounds out the evaluation, with results described as strong even for low-resource languages. Model weights are noted as coming soon to Hugging Face, requiring a multi-GPU cluster to run locally.


📺 Source: Fahd Mirza · Published June 30, 2026
🏷️ Format: Review

1 Item

Channels