• Mobile Cloud and its partners have combined domestic GPUs and neuromorphic chips in an AI inference system that divides model calculations between the two processors
  • The developers report more than 40% lower operating costs than a comparable domestic GPU cluster in tests using DeepSeek V4 Flash

The fact

China Mobile's Mobile Cloud business and its partners have launched an AI inference system that combines domestic GPUs with neuromorphic chips. The system was presented during the China Computing Conference in Langfang, Hebei, held from 11 to 13 September.

The project was developed with the CETC Nanhu Research Institute, Lynxi Technology, Iluvatar CoreX, Tsinghua University and Peking University. The developers said tests using DeepSeek V4 Flash delivered more than twice the price-performance and cut operating costs by more than 40% compared with a similar domestic GPU cluster. The figures come from the development team and have not been independently verified.

The system splits work between the two chip types. Attention calculations run on domestic GPUs, while FFN work is assigned to neuromorphic processors. A model compiler, high-speed interconnect protocol and common inference engine manage the division of tasks, scheduling and results. Iluvatar separately reported 654 million yuan of inference-product revenue in the first half of 2026, equal to 69.2% of its total revenue.

The assessment

The idea is simple: do not make one type of processor handle every part of the model. The GPU keeps the work that suits a general-purpose accelerator, while the neuromorphic chip takes calculations the developers believe it can run more efficiently. If that division holds up in production, a cloud operator could lower the cost of serving AI requests without replacing its entire GPU stack.

The difficult part is everything between the two chips. Data has to move at the right time, the software has to decide where each calculation runs, and neither processor can spend too long waiting for the other. A saving measured on one DeepSeek configuration may shrink when requests are longer, traffic is heavier or the model changes.

For BTW readers, the 40% claim therefore needs a whole-system comparison. Customers would need to run the same workload at the same response time and output quality, while counting both processors, the interconnect and the software. If the mixed system still costs less under those conditions, the architecture becomes commercially interesting.

What to watch

Customer trials should show complete-system cost, response time and throughput rather than another peak benchmark. Results across several models and traffic patterns would show whether the reported saving survives outside the DeepSeek V4 Flash test. Commercial orders or deployments by cloud and enterprise customers would provide stronger evidence than further demonstrations.