A Chinese lab opens the software layer beneath its AI chips

Six open-source components aim at the programming stack that keeps rival accelerators dominant.

ChipNews Staff
2 Min Read

The hardest part of replacing an accelerator is rarely the transistor. It is everything above it, which is why a Chinese AI lab and Huawei have published a set of open-source components aimed at the software that sits directly on the silicon.

The lab posted the projects to its WeChat account on September 30, and Reuters confirmed the collaboration the same day. The releases span kernel authoring, matrix math, inter-chip communication, attention routines and data selection.

Three of them carry most of the weight. One provides a portable language for writing AI kernels. Another handles collective traffic across a cluster. A third packages the operations needed by the lab’s own attention design.

Engineers previously had to issue device-specific commands to get performance out of the hardware. The new layer lets them express a computation once and leave the mapping to the compiler, which is the same bargain that made rival toolchains sticky over the past decade.

The lab frames that as a precondition rather than a convenience. Without a general language that is both easy to program and able to reach the hardware’s limits, it argues, a domestic ecosystem cannot stand on its own.

The two firms also showed a supernode that ties 128 accelerators into one logical machine.

What matters is the breadth. A faster chip can be copied; the libraries, compilers and primitives around it take years to reproduce, and they are the reason incumbents keep their customers.

The lab has been moving workloads onto Huawei silicon through 2026 as export controls block the most advanced Nvidia parts. Its newest model was tuned for domestic hardware, and a planned Inner Mongolia site is expected to hold more than 160,000 accelerators.

Share This Article