Moore Threads Unveils MTT C256 Supernode: Turning 256 GPUs into a Single Supercomputer at WAIC 2026
At WAIC 2026, Moore Threads showcased its MTT C256 supernode, a groundbreaking architecture that interconnects 256 GPUs via a single-layer Scale-up network, transforming them into a unified data-center-class supercomputer.
Shanghai, July 20, 2026 – The 2026 World Artificial Intelligence Conference (WAIC), which kicked off on July 17, has become a stage for Moore Threads to unveil its latest breakthrough in high-performance computing. Under the theme "Token Era, Everything Intelligent," the company demonstrated three AI factories—Model Training Factory, Token Production Factory, and Agent Production Factory—alongside the debut of the MTT C256 supernode, a system that aggregates 256 GPUs into one logical supercomputer.
Traditional data center networks struggle with bandwidth loss and high latency as clusters scale to thousands of GPUs. Moore Threads tackles this with a novel one-layer Scale-up network that for the first time enables full interconnect among 256 GPUs within a single network hop. Previously, the industry limit for such direct connectivity was 64 cards.
Key Advantages of MTT C256
- Extreme Interconnect: The single-layer architecture reduces card-to-card communication latency to sub-microsecond levels and dramatically improves bandwidth utilization, enabling more efficient distributed parallel training of large models.
- High-Density Deployment: The entire 256-GPU cluster fits into just two standard racks, boosting GPU density per rack while saving valuable data center floor space.
- Ecosystem & Engineering Compatibility: Using a standard 2U node design, the MTT C256 fully integrates with existing AI software and hardware ecosystems, allowing seamless scaling from 10,000 to 100,000 GPUs.
"National Chip Trains National Model" – Live Demonstrations
Moore Threads also showcased real-world results from its KUAE computing cluster powered by domestic GPUs:
- Trillion-Parameter MoE Model Training: Using a 10,000-GPU KUAE cluster and the KUAE Training Suite, Moore Threads completed full training of the 236B-parameter MoE model from scratch with over 25T tokens. The cluster achieved more than 90% effective training time, with a loss curve matching that of leading international GPUs.
- 5D World Model Native Training: In collaboration with Peking University’s EvoPhys team, the system supported the first full-stack native training of the global 5D world model EvoPhys-World, relying entirely on the MUSA software stack—validating scalability for complex visual generation and physical simulation.
- Embodied AI Brain Training: Partnering with Beijing Academy of Artificial Intelligence (BAAI) under the FlagOS-Robo framework, Moore Threads completed end-to-end training of the RoboBrain 2.5 general-purpose embodied brain model on its MTT S5000 thousand-card cluster, proving the viability of domestic computing for embodied intelligence workloads.
The MTT C256 supernode represents a leap in GPU clustering efficiency and signals China's push toward self-reliant, large-scale AI infrastructure.