Announcement of the Cerebras processor ― Cerebras Wafer Scale Engine (WSE) or the Cerebras silicon wafer scale engine ― at the annual Hot Chips 31 conference. Looking at this silicon monster, it's not even surprising that it was managed to be brought to life. What amazes is the boldness of the concept and the efforts of the developers who dared to create a chip with an area of 46,225 square millimeters with sides of 21.5 cm. It takes an entire 300 mm wafer to manufacture one processor. With the slightest mistake, the defect rate reaches 100%, and the price of the issue is hard to imagine.

The production of the Cerebras WSE is handled by TSMC. The process technology is 16 nm FinFET. This Taiwanese manufacturer has also earned a monument for the release of Cerebras. Producing such a chip required the highest craftsmanship and the solution to numerous problems, but it was worth it, the developers assure. The Cerebras chip is essentially a supercomputer on a chip with incredible bandwidth, minimal power consumption, and fantastic parallelism. Currently, it is the ideal solution for machine learning, enabling researchers to tackle extremely complex problems.

Each Cerebras WSE chip contains 1.2 trillion transistors, organized into 400,000 AI-optimized computing cores and 18 GB of local distributed SRAM memory. All of this is connected via a mesh network with an overall performance of 100 petabits per second. Memory bandwidth reaches 9 PB/s. The memory hierarchy is single-level. There is no cache memory, no overlaps, and access latencies are minimal. This is the perfect architecture for accelerating AI-related tasks. Hard numbers: compared to the most modern graphical cores, the Cerebras chip provides 3000 times more memory on the chip and 10,000 times greater memory bandwidth.

Cerebras computational cores — SLAC (Sparse Linear Algebra Cores) — are fully programmable and can be optimized for use with any neural networks. Moreover, the core architecture initially filters data represented by zeros. This frees computational resources from the need to perform redundant multiplication operations with zero, which for sparse data workloads means accelerated calculations and maximum energy efficiency. Thus, the Cerebras processor is hundreds or even thousands of times more efficient for machine learning in terms of chip area used and power consumption compared to current AI and machine learning solutions.

Manufacturing a chip of this size a host of unique solutions. It had to be packaged almost by hand. There were challenges with delivering power to the chip and its cooling. Heat dissipation was only possible using liquid cooling and required zoned delivery with vertical circulation. Nevertheless, all issues were resolved, and the chip became operational. It will be interesting to learn about its practical applications.

Source: 3dnews.ru
