ByteDance trains 10-trillion-parameter AI model, tripling China's largest

ByteDance has begun pre-training a language model with roughly 10 trillion parameters. The project aims to close the gap with leading global systems such as Anthropic's Mythos and signals the Chinese company's determination to compete at the top tier of the industry, despite restricted access to advanced chips.
The scale alone makes a statement. For context, Moonshot's Kimi K3, one of China's largest models, has about 2.8 trillion parameters. ByteDance's new effort is more than three times that size.
Parameter count, however, does not guarantee superiority. The industry has learned that data quality, training methodology, and computational efficiency often matter as much as raw scale. Still, committing compute resources to a model of this size is a declaration in itself — an intent to compete at the frontier rather than ship a capable also-ran.
The project is in its early stages, and the final performance will depend on multiple factors beyond sheer size. Yet the very existence of such training underscores how aggressively Chinese firms are pushing frontier AI under hardware sanctions.


