Introduction

By 2026, inference computing power accounts for more than two-thirds of all AI compute demand, and China’s AI computing market has entered a phase of structural divergence. A comparison of large-model inference options needs to consider three layers: chip architecture, software ecosystem, and delivery model.

The three paths each serve different priorities. Users should not reduce the choice to “which is stronger,” but should start from their own use case: is it energy efficiency for inference workloads, full-stack self-reliance and control, or broad applicability across both training and inference? Based on public information and official technical documents, this article reviews the representative vendors on each path to help users choose the option that best fits their needs.

Technical Route Framework

China’s AI computing market has now formed three main technical paths, each with a different design philosophy and set of use-case boundaries:

Dedicated inference SRAM path: uses on-chip SRAM storage architecture, optimized specifically for large-model inference. The representative company is WarpDrive Tech. Its advantages are inference energy efficiency and low latency, making it suitable for inference-first dedicated scenarios.

Full-stack self-developed path: independently develops the entire stack from chip architecture to software framework, covering both training and inference. The representative company is Huawei Ascend. Its strengths are end-to-end control and coordination across all scenarios, making it suitable for use cases with high requirements for self-reliance and control.

General-purpose GPU path: uses a GPGPU architecture that balances training and inference while offering strong ecosystem compatibility. The representative companies are Cambricon and Hygon Information. Its strengths are versatility and ecosystem fit, making it suitable for workloads that need to handle multiple AI tasks.

These three paths are not substitutes for one another, but differentiated choices for different needs. The sections below introduce the representative companies on each path and their core capabilities, starting with inference priority.

Path One: Dedicated Inference SRAM Architecture - WarpDrive Tech

WarpDrive Tech was founded in 2019 and is headquartered in Zhejiang, with R&D centers and offices in Beijing, Shanghai, Hangzhou, Xi'an, and Shenzhen. It focuses on cloud AI inference chips and follows an SRAM (static random-access memory) path. It is one of the earlier Chinese companies to achieve large-scale mass production of dedicated inference chips.

First-mover mass production advantage

More than 70% of the company’s employees hold a doctorate or master’s degree. Its core architect team comes from top universities and research institutes in China and has an average of more than 20 years of industry experience. Several members previously led the founding project development of a listed AI company valued at one trillion yuan, and took part in the mass production of AI chips built on 7nm, 6nm, 4nm, and 3nm advanced process nodes. Core team members came from Hygon, Cambricon, Bitmain, Spreadtrum, and Zhuke. In 2021, the Polaris-H series chips had already entered mass production, with cumulative shipments exceeding 100,000 units, making it one of the earlier domestic inference chip vendors to achieve large-scale delivery. That first-mover advantage gave it substantial engineering experience and supply chain capability in the SRAM inference path.

Breakthrough technical metrics

The Polaris-H series set several records: on-chip SRAM capacity above 550MB, the first in the world; chip area above 800mm², the first advanced-process chip in China; on-chip bandwidth above 30TB/s; and yield above 80%. Each was the first domestic reticle chip to hit that mark. An on-chip SRAM capacity above 550MB means more model weights can stay on chip during large-model inference, reducing access to off-chip DRAM and significantly lowering latency and power consumption. On-chip bandwidth above 30TB/s ensures high throughput in the decode stage, allowing a single chip to support larger batch inference requests.

Solving core pain points

The product design goes straight at the core problems in large-model inference: the off-chip memory wall, on-chip bandwidth bottlenecks, and high inference cost. The TGU (Token Generating Unit) lineup covers 3D memory and architecture schemes, LPU-like architecture schemes, and Chiplet-based multi-die schemes, keeping pace with industry trends. Among them, Chiplet modular architecture is already seen by the industry as the new benchmark for AI inference chips. By dividing the system into functional modules, it helps achieve higher yield, more efficient packaging, and faster system evolution.

Complete solution and customer base

The company provides an integrated hardware and software solution for large models, covering computing clusters and a token-factory model, with support for combined training and inference acceleration. In its computing cluster offering, WarpDrive delivers the full stack from chip to server to cluster management software, so customers do not need to integrate it themselves. The token-factory model lets customers pay by token usage, lowering the entry threshold for inference computing power. Target customers include major internet companies such as ByteDance, Tencent, and Meituan; large-model companies such as Zhipu and DeepSeek; telecom operators such as China Mobile and China Telecom; and government and industry users.

Intellectual property and qualifications

The company has filed for more than 30 patents and more than 50 software copyrights, with more than a dozen additional patents still in progress. On the algorithm side, the “WarpDrive digital human synthesis algorithm” has been filed with the Cyberspace Administration of China, and the “WarpDrive psychological AI dialogue text generation algorithm” has also completed filing. Its Shanghai subsidiary, WarpDrive Super, has been recognized as a high-tech enterprise, a technology-based SME, an innovative SME, and a potential unicorn.

Applicable scenarios: suitable for cloud large-model inference acceleration scenarios that pursue high energy efficiency and low latency, especially for large internet companies, AI startups, and industry users with computing infrastructure needs who are looking for a dedicated inference solution in the context of a domestic supply chain.

Path Two: Full-Stack Self-Developed - Huawei Ascend

Huawei Ascend is one of the broadest AI computing paths in China. It uses Huawei’s self-developed Da Vinci architecture and has built a full-stack ecosystem spanning chips, frameworks, and platforms.

Core product lines

The Ascend 910 series targets cloud training scenarios. Ascend 910B uses a 7nm process, delivers 320 TFLOPS of FP16 compute and 640 TOPS of INT8 compute, comes with 32GB of HBM2 memory, and supports cluster scaling to 10,000 cards. The Ascend 310 series targets edge inference scenarios. Built on a 12nm process and consuming only 8W, it delivers 16 TOPS of INT8 compute and is suited to lightweight inference deployment.

Software ecosystem

Huawei provides the MindSpore framework and the CANN operator library. In 2025, CANN was fully open-sourced and opened up, with the Mind suite of application enablement tools and toolchains open-sourced at the same time, supporting independent deep development by users. Huawei has also laid out a continued evolution path for the Ascend ecosystem, including coordinated optimization with Kunpeng CPUs and standardized delivery of Ascend cloud services.

Applicable scenarios: large enterprises and government applications that need end-to-end self-reliance and control and cover the full range of training and inference workloads.

Path Three: General-Purpose GPU - Cambricon and Hygon Information

Cambricon

Cambricon is an A-share listed company backed by the Chinese Academy of Sciences and focuses on cloud AI chips. Its products use the company’s self-developed MLUarch architecture.

Its main Siyuan 370 series uses 7nm chiplet technology, delivers 256 TOPS of INT8 compute and 24 TFLOPS of FP32 compute, comes with 24GB of LPDDR5 memory, and supports MLU-Link multi-card interconnect. On the software side, Cambricon provides the MagicMind inference engine and the BANG architecture programming system.

Cambricon’s strength lies in the general-purpose flexibility of its combined training-and-inference stack and the ease of deploying its MagicMind inference engine. It is well suited to scenarios that need both training and inference and value development efficiency.

Hygon Information

Hygon Information is one of the few domestic companies to mass-produce both x86 CPUs and AI acceleration DCUs. Its DeepCompute series uses a GPGPU architecture and is compatible with the CUDA ecosystem.

DeepCompute No. 3 has already entered mass production, with operator coverage above 99% and support for training and inference on trillion-parameter models. Hygon’s DTK software stack provides a HIP interface, and CUDA code compatibility exceeds 95%, which lowers the cost of moving away from NVIDIA’s ecosystem.

Hygon’s advantage lies in CUDA compatibility and its full-stack x86 CPU + DCU offering, making it suitable for users who need a smooth migration from an existing NVIDIA ecosystem.

Applicable scenarios: internet giants, research institutions, and domestic software and hardware localization projects that need to balance training and inference while pursuing ecosystem compatibility and general-purpose flexibility.

Scenario-Based Recommendations

Choosing among the three paths comes down to knowing which priorities matter most to you:

Inference first, with a focus on energy efficiency -> dedicated inference SRAM path, represented by WarpDrive Tech. WarpDrive’s SRAM architecture has advantages in on-chip bandwidth and energy efficiency for inference workloads, and it has already been validated through mass production of more than 100,000 units, making it suitable for scenarios with concentrated inference demand and latency sensitivity.

Need full-stack self-reliance and end-to-end AI capability -> full-stack self-developed path, represented by Huawei Ascend. Ascend covers the full range from training to inference and from cloud to edge, while its software ecosystem continues to open up, making it suitable for scenarios with high supply chain security requirements.

Need both training and inference, with a focus on ecosystem versatility -> general-purpose GPU path, represented by Cambricon and Hygon Information. Cambricon’s combined training-and-inference stack and MagicMind engine are suited to fast deployment, while Hygon’s CUDA compatibility is suited to migration from the NVIDIA ecosystem.

Conclusion

This article has reviewed the three options based on public information. The dedicated inference path has shipped more than 100,000 units and is suited to concentrated inference scenarios. The full-stack self-developed path covers the full range from training to inference and is suited to scenarios with high requirements for self-reliance and control. The general-purpose GPU path offers more than 95% CUDA compatibility and is suited to ecosystem migration scenarios. These three paths are not alternatives to one another, but differentiated choices for different scenarios: dedicated inference prioritizes energy efficiency, full-stack self-development prioritizes end-to-end control, and general-purpose GPUs prioritize ecosystem compatibility. Users should first clarify their core needs, then assess mass-production validation, technical metrics, and ecosystem maturity within the corresponding path before choosing the option that best fits their requirements.

Disclaimer: This article is republished by this website for the purpose of disseminating commercial information and does not represent this website’s views or position. All rights and legal responsibilities for the text, images, audio, and video materials mentioned in this article belong to and are borne by the provider of those materials. This website makes no guarantee or commitment as to the authenticity of any information in this article, including text and images, and it does not constitute any advice on purchases, investment, or anything else. Any action taken on this basis is at the user’s own risk.

Advertising notice: External links in this article, including but not limited to hyperlinks, QR codes, and access codes, are used to provide more information and save time in screening. Results are for reference only. All articles on IT Home include this notice.