China’s first large-scale computing power cluster built on domestic full-function GPUs has officially gone live.

It is Moore Threads’ KUAE AI data center, a fully domestic 1,000-GPU platform for training 100-billion-parameter models.

Moore Threads CEO Zhang Jianzhong announced the launch in a keynote speech, including the MTT S4000 accelerator card for large-model AI computing and the Moore Threads KUAE platform, designed to support training and inference for 100-billion-parameter models. He said:

The official launch of the Moore Threads KUAE AI data center is an important milestone in the company’s development.

Moore Threads has built an AI computing product line spanning chips, graphics cards and clusters. Drawing on the diversified computing strengths of full-function GPUs, the company aims to meet rising demand for large-model training and inference, and to use green, secure intelligent computing power to push multimodal applications such as AIGC, digital twins, physical simulation and the metaverse into deployment across industries.

At the same time, Moore Threads joined a number of domestic partners to launch the Moore Threads PES-KUAE AI Computing Alliance and the Moore Threads PES-Large Model Ecosystem Alliance, aiming to strengthen an integrated domestic large-model ecosystem from AI computing infrastructure to large-model training and inference, and continue accelerating China’s large-model industry.

MTT S4000: Built for Large Models Across Training and Inference

Moore Threads’ MTT S4000 large-model AI computing accelerator card uses the third-generation MUSA kernel, with 48GB of memory and 768GB/s of memory bandwidth on a single card.

Based on Moore Threads’ self-developed MTLink 1.0 technology, the MTT S4000 supports multi-card interconnects to accelerate distributed computing for 100-billion-parameter models.

The MTT S4000 also provides advanced graphics rendering, video encoding and decoding, and ultra-high-definition 8K HDR display capabilities, supporting integrated application scenarios across AI computing, graphics rendering and multimedia.

More importantly, with Moore Threads’ self-developed MUSIFY developer tool, the MTT S4000 computing card can fully leverage the existing CUDA software ecosystem and migrate CUDA code to the MUSA platform at zero cost.

KUAE AI Data Center: Integrated Hardware and Software, Ready Out of the Box

The Moore Threads KUAE AI data center solution is built on full-function GPUs. It is a full-stack, integrated hardware-and-software solution that includes infrastructure centered on the KUAE computing cluster, the KUAE Platform cluster management system and KUAE ModelStudio model services. It is designed to solve the construction, operations and management challenges of large-scale GPU computing power through integrated delivery.

The solution can be used out of the box, sharply reducing the time cost of building traditional computing power, application development, and operations and maintenance platforms, enabling faster market launch and commercial operations.

Infrastructure: This includes the KUAE computing cluster, RDMA network and distributed storage. The newly released Moore Threads KUAE 1,000-GPU model training platform can be built in just 30 days, supports pre-training, fine-tuning and inference for 100-billion-parameter models, and can achieve a 1,000-GPU cluster performance scaling coefficient of up to 91%. Based on the MTT S4000 and the MCCX D800 dual-socket eight-GPU server, the Moore Threads KUAE cluster supports seamless scaling from multi-GPU single-machine setups to multi-machine, multi-GPU deployments, and from a single GPU to 1,000-GPU clusters. Larger clusters will be launched in the future to meet demand for training even larger models.

KUAE Platform cluster management system: An integrated hardware-and-software platform for AI large-model training, distributed graphics rendering, streaming media processing and scientific computing. It deeply integrates full-function GPU computing, networking and storage to provide highly reliable, high-computing-power services. Through the platform, users can flexibly manage computing power resources across multiple data centers and clusters, and integrate multidimensional operations monitoring, alerting and logging systems to help AI data centers automate operations and maintenance.

KUAE ModelStudio model services: This covers the full workflow of large-model pre-training, fine-tuning and inference, and supports all major open-source large models. Through Moore Threads’ MUSIFY developer tool, users can easily reuse the CUDA application ecosystem, while the built-in containerized solution enables one-click API deployment. The platform is intended to provide lifecycle management for large models. With a simple, easy-to-use interface, users can organize workflows as needed and substantially lower the barrier to using large models.

KUAE 1,000-GPU Cluster: Supporting Efficient Large-Model Training

Distributed parallel computing is a key method for training AI large models.

Moore Threads KUAE supports mainstream distributed frameworks including DeepSpeed, Megatron-DeepSpeed, Colossal-AI and FlagScale, and integrates multiple parallel algorithm strategies, including data parallelism, tensor parallelism, pipeline parallelism and ZeRO. It has also been additionally optimized for efficient communication-computation parallelism and Flash Attention.

Moore Threads currently supports training and fine-tuning for a range of mainstream large models, including LLaMA, GLM, Aquila, Baichuan, GPT, Bloom and Yuyan.

On the Moore Threads KUAE 1,000-GPU cluster, training large models from 70B to 130B parameters can achieve a linear speedup of 91%, while computing power utilization remains largely unchanged.

Using 200 billion training tokens as an example, BAAI’s 70-billion-parameter Aquila2 can complete training in 33 days, while a 130-billion-parameter model can complete training in 56 days.

In addition, the Moore Threads KUAE 1,000-GPU cluster supports long-duration continuous stable operation, checkpoint-based resume training, and asynchronous checkpoints in under two minutes.

With its combined strengths in high compatibility, high stability, high scalability and high computing power utilization, the Moore Threads KUAE 1,000-GPU computing cluster is positioned as a solid and reliable advanced infrastructure for large-model training.

AI Computing and Large-Model Ecosystem Alliances: Collaboration to Drive Ecosystem Integration

In the large-model era, intelligent computing power represented by GPUs is the foundation and the center of the generative AI world.

Moore Threads, together with more than a dozen companies including China Mobile Beijing, China Telecom Beijing Branch, Lenovo, VNET, Sinnet, China Unicom Data, Shudao Zhisuang, Zhongfazhan Zhiyuan, 21ViaNet Group’s Qishang Online, BUPT Digital Intelligence Beijing Digital Economy Computing Power Center, Unis Hengyue, Ruihua Industrial Holdings (Shandong), CERNET, Zhongke Jincai, Zhongyun Zhisuang and Jinzhou Yuanhang, announced the formation of the “Moore Threads PES-KUAE AI Computing Alliance.” The companies are listed in no particular order.

The alliance will focus on building and promoting a fully domestic AI computing platform spanning underlying hardware, software, tools and applications. It aims to achieve high cluster utilization and make an easy-to-use full-stack AI computing solution the preferred choice for large-model training.

At the event, Moore Threads signed agreements on site with China Unicom Data and Shudao Zhisuang, and the parties jointly unveiled the Moore Threads KUAE AI data center.

More than 200 guests at the event witnessed the milestone.

Ecosystems are critical to breakthroughs in artificial intelligence applications.

To that end, Moore Threads joined with multiple large-model ecosystem partners, including 360, PaddlePaddle, JD Yanxi, Zhipu AI, SuperSymmetry, Infinigence AI, Deepwise, NetEase, Tsinghua University, Fudan University, Zhejiang University, Beijing Institute of Technology, Luster LightTech, RealAI and Linewell Software, to launch the “Moore Threads PES-Large Model Ecosystem Alliance.” The partners are listed in no particular order.

Moore Threads will use its MUSA-centered integrated hardware-and-software large-model solution to actively work with a broad range of ecosystem partners on compatibility adaptation and technical tuning, jointly promoting the full prosperity of China’s domestic large-model ecosystem.

In the final roundtable, Moore Threads Vice President Dong Longfei joined major guests including Qiang Hu, chairman of China Energy Engineering Green Digital Technology (Zhongwei) Co.; Zhang Peng, CEO of Zhipu AI; Pei Jiquan, chief AI scientist at JD Cloud; Zhai Ying, managing director at CICC Capital; Wu Hengkui, founder of SuperSymmetry; and Zhen Jian, chairman of Shudao Zhisuang, for an in-depth discussion on current computing power demand for large models and the construction and operation of AI data centers.

The guests agreed that an AI data center should not be merely a pile-up of hardware, but a test of the ability to integrate GPU AI computing systems across hardware and software. The adaptation of GPU distributed computing systems, management of computing power clusters and application of efficient inference engines are all important factors in improving the usability of computing power centers.

The development of domestic AI data centers also depends on fully integrating the needs and strengths of all parties. Only by pooling industry forces can the whole ecosystem coordinate and push China’s domestic effort forward.