To be included in the upcoming branch of the Linux kernel 6.13, a patch has been proposed with a revised implementation of the CRC32C checksum algorithm. The code for the CRC32C implementation has been reduced by approximately 10 times (from 4546 to 418 bytes). With retpoline protection against Spectre-class attacks disabled, the performance gain using the new implementation reaches 11.8% on AMD Zen 2 processors, 6.4% on Intel Emerald Rapids, and 4.8% on Intel Haswell. With retpoline enabled, the performance improvement is more pronounced, reaching 66.8% on systems with Intel Emerald Rapids processors, 35.0% on Intel Haswell, and 29.5% on AMD Zen 2. retpoline enabled | 512 | 833 | 1024 | 2000 | 3173 | 4096 | βββββββ+ββ-+ββ-+ββ-+ββ +ββ-+ββ-+ Intel Haswell | 35.0% | 20.7% | 17.8% | 9.7% | -0.2% | 4.4% | Intel Emerald Rapids | 66.8% | 45.2% | 36.3% | 19.3% | 0.0% | 5.4% | AMD Zen 2 | 29.5% | 17.2% | 13.5% | 8.6% | -0.5% | 2.8% | retpoline disabled: | 512 | 833 | 1024 | 2000 | 3173 | 4096 | βββββββ+ββ-+ββ-+ββ-+ββ +ββ-+ββ-+ Intel Haswell | 3.3% | 4.8% | 4.5% | 0.9% | -2.9% | 0.3% | Intel Emerald Rapids | 7.5% | 6.4% | 5.2% | 2.3% | -0.0% | 0.6% | AMD Zen 2 | 11.8% | 1.4% | 0.2% | 1.3% | -0.9% | -0.2%
The original version of CRC32C included 128 unrolled cycles, which resulted in quite a large code base. Since modern processors with support for out-of-order execution can execute commands in parallel, such optimization of loop branch instructions became redundant and only resulted in excessively large code. Instead of 128 iterations, the new version retains only 4, which not only significantly reduces the code size but also speeds up the operation.
Source: opennet.ru
