After two months of development, Linus Torvalds has released the Linux kernel 7.2. Among the most notable changes are: the USB4STREAM data streaming mechanism, performance optimizations for btrfs, xfs, and ext4, continued removal of code for i486 CPU support, the ability to create nested schedulers SCHED_EXT, reduced memory usage in the swap subsystem, accelerated unnamed pipes, support for Intel MBEC and AMD GMET extensions in KVM, removal of the AppleTalk protocol, and initial support for HDMI 2.1 FRL in the AMDGPU driver.
Key innovations in kernel 7.2 (1, 2, 3):
- Disk subsystem, input/output, and file systems
- In the iomap mechanism, the unnecessary memset function call for already completed iterations in iomap_iter() has been removed, which increased input/output operations per second (IOPS) by 5% during tests when using the ext4 and xfs file systems on high-intensity I/O with fast NVMe drives.
- In XFS, stable support for zoned storage devices has been declared stable (partitioning into zones of block groups or sectors that allow only sequential addition of data with full group updates).
- In Btrfs, support for large memory page folios has been enabled by default, allowing for reduced overhead and improved performance during intensive sequential I/O. Experimental support for huge folios, up to 2 MB in size, has been added. A new ioctl GET_CSUMS has been introduced to obtain checksum information in user space, for example, for the mkfs utility and deduplication optimization. Sequential data write performance has increased by 15% and direct I/O by 59%.
- In the Ext4 file system, the implementation of the fast commit mechanism has been significantly redesigned to eliminate competing and deadlock situations. The export of inode snapshot statistics has been added through /proc/fs/ext4/*/fc_info. Directory hash calculation performance has been optimized (for names of 255 characters, nearly double speed, 64 characters 27%, 32 characters — 11%).
- In F2FS, support for returning fserror errors has been added, allowing tracking of file system issues from user space. The time spent in the context of interrupt handling has been reduced.
- The Device Mapper (DM) has added a new handler called dm-inlinecrypt for organizing transparent encryption and decryption of block devices, using hardware devices with inline encryption capabilities.
- Documentation has been proposed for adding new file systems to the kernel.
- In NFS, the default block size has been increased to 4 MB on systems with at least 16 GB of RAM (for manual adjustment of the block size, use /proc/fs/nfsd/max_block_size). Support for directory delegation has been added, allowing certain operations on that directory to be performed without state change checks. server.
- The SMB server now supports files stored in compressed form, as well as data compression during network transmission.
- The new implementation of NTFS (ntfsplus) has added support for Windows symbolic links and ensures proper handling of many types of metadata corruption.
- The fscache backend for caching data in EROFS (Enhanced Read-Only File System) has been removed, which was deprecated two years ago.
- The Ceph file system has added support for manually resetting client sessions.
- The 9P file system has been optimized for better performance in scenarios like project builds.
- Flags have been added to the file_getattr() system call to retrieve information about case sensitivity in the file system. The FS_XFLAG_CASEFOLD flag indicates that filename checks are case insensitive, while the FS_XFLAG_CASENONPRESERVING flag means that case information is not preserved when creating new filenames. These flags can be applied in NFS clients that ignore case sensitivity.
- An O_EMPTYPATH flag has been added to the openat2() system call, allowing the passing of an empty file path. In this case, the path to the file being opened is determined based on the provided file descriptor.
- An OPENAT2_REGULAR flag has been added to the openat2() system call, allowing the opening of only regular files (attempting to open a special file, such as a socket, pipe, or device, will return an EFTYPE error).
- Memory and system services
- The USB4STREAM mechanism has been implemented for data streaming between computers connected via USB4 ports. A device /dev/tbstreamX has been added, which allows reading and writing data using standard read() and write() functions, similar to reading and writing to files. For example, information can be sent from one host using the command "echo hello > /dev/tbstream0", while on the other side, it can be read with "cat /dev/tbstream0". The USB4STREAM mechanism can be combined with the ability to establish a network connection over USB4 cable (thunderbolt_net) or used separately as needed for data transfer between applications that do not support network sockets.
- The second series of changes has been included to discontinue support for i486 processors. More than 13 lines of code related to the floating-point emulation block for processors without FPU have been removed. Support for i486 processors without hardware CX8 (compare and exchange 8 bytes) and TSC (CPU cycle counter used in the task scheduler) operations has been removed, and the emulation code for them has been eliminated.
- The orphaned support for AMD Geode processors, used in the OLPC XO-1 computer, has been declared.
- Support for updating the Intel TDX (Trusted Domain Extensions) mechanism implementation has been added, which is used for encrypting the memory of guest systems. TDX is implemented as a special software runtime module that is transferred from BIOS from Flash memory to RAM during initial boot. The kernel has been enhanced with capabilities to manage this module and replace it with a newer version on the running system without requiring a reboot.
- A new GPU resource allocation scheduler (Fair GPU scheduler) has been implemented to determine the order of execution of GPU tasks sent by processes utilizing the GPU. Instead of using the traditional FIFO request queue for the GPU, the new scheduler employs fair resource allocation mechanisms, referencing the CFS (Completely Fair Scheduler) task scheduler that uses a time-sharing model for transitioning to the next process. The most noticeable effect of using the new scheduler is observed during the parallel execution of interactive tasks that actively engage the GPU. At the last moment before the release of kernel 7.2, the inclusion of the Fair GPU scheduler was canceled, and the old FIFO scheduler was reinstated due to the need for debugging regressions that led to performance drops and 100% GPU load when launching certain games in Proton.
- Changes in the eBPF subsystem: The ability to attach one BPF program to multiple tracepoints has been added. BPF programs attached to tracepoints now have access to the memory of components operating in user space, with proper handling of memory access faults (page fault). The bpf() system call has been enhanced to support standard attributes (log_buf, log_size, log_level, and log_true_size), enabling a unified transfer of metadata across all BPF commands, not limited to BPF_PROG_LOAD, BPF_BTF_LOAD, and BPF_MAP_CREATE commands. The limit on passing more than 5 parameters in BPF functions has been removed. Safe access to shared memory bpf_arena has been added without the risk of encountering memory access faults (page fault). A new variant of the BPF hash map structure allowing dynamic resizing has been implemented.
- The output from "/proc/interrupts" regarding interrupt statistics has been optimized, alongside updated structures for storing interrupt counters and added caching.
- The generation of the file "/proc/filesystems", used in libselinux, has been accelerated.
- The task scheduler now supports load balancing between CPU cores, taking into account the state of the processor's internal cache. The scheduler now attempts to group processes that use shared resources, such as threads from a single process, for use in the context of the same upper-level cache, which enhances data access efficiency by increasing the likelihood of finding the necessary data in the cache.
- In the SCHED_EXT mechanism, which allows the use of BPF to create CPU schedulers, the implementation of the capability for creating nested schedulers (sub-schedulers) has continued, enabling dedicated task schedulers for each cgroup.
- The transfer of changes from the Rust-for-Linux branch continues, related to the use of the Rust language as a second language for driver and kernel module development (Rust support is not enabled by default and does not include Rust in the list of mandatory kernel build dependencies). Rust's capability in the kernel has been implemented for the s390 architecture. The package 'zerocopy' with fast memory manipulation primitives for 'unsafe' code has been included.
- The minimum version of the LLVM toolchain required for building the kernel has been raised to 17.0.1.
- In the minimalistic C library nolibc, included with the Linux kernel source code and providing a wrapper over basic system calls, support for OpenRISC and 32-bit PA-RISC architectures has been implemented.
- Optimizations have been made in the swap subsystem that improve performance and reduce memory usage within the subsystem by eliminating overhead when storing static metadata and unifying the operation with anonymous and shared memory when using volumes. Memory consumption decreases significantly; for instance, when mounting a 1 TB swap partition, memory usage is reduced by approximately 512 MB.
- The efficiency of the memory eviction mechanism has been improved, removing or transferring memory areas to swap space to free up memory when it is scarce in the system. In some workloads, such as testing MongoDB with YCSB (Yahoo! Cloud Serving Benchmark), performance gains of up to 30% have been observed.
- A new command "make sbom" has been added to the build system to generate SBOM (Software Bill Of Materials) lists, reflecting the components, libraries, and dependencies used in the current kernel build, as well as licensing information obtained from SPDX headers in code files.
- The implementation of unnamed pipes has been optimized to improve lock handling (memory allocation operations have been moved outside the lock scope), increasing the throughput of unnamed pipes by 21-48% and reducing latency by 17-33%.
- Virtualization and Security
- The slab allocator has been enhanced with the ability to use allocation tokens implemented in Clang 22. Tokens allow the marking of memory allocation operations with unique identifiers and organize separate placement of different object types to complicate the exploitation of vulnerabilities caused by buffer overflow (with separation, exploiting a buffer overflow in one type of object does not easily affect other types).
- The AF_ALG mechanism, exploited in the Copy Fail vulnerability for modifying data in the page cache, has been deprecated and is slated for removal in a future release. AF_ALG allows the use of hardware accelerators for cryptographic computations in the kernel's Crypto API, but is used in rather specific situations. In kernel 7.2, support for asynchronous I/O, old drivers, and the zero-copy mechanism in the skcipher and aead implementation has been removed from AF_ALG. Only software implementations of cryptographic algorithms remain, and support for hardware crypto accelerators in the kernel's Crypto API has been removed, as AF_ALG significantly expands the attack surface without providing performance gains compared to the cryptographic implementation in user space. AF_ALG was used in the Cryptsetup toolkit, but its support was removed in the recent release 2.8.7.
- The IMA (Integrity Measurement Architecture) mechanism, which allows an external service to verify the state of kernel subsystems to ensure their authenticity, has been updated to support exporting internal tables with measurement results to user space, removing them from kernel buffers to save memory.
- The Landlock module, which provides unprivileged programs with means to restrict the use of Linux kernel objects (file hierarchies, network sockets, ioctl, etc.), has added support for access control to UDP sockets, as well as the ability to selectively disable the logging of object blocking information to prevent the log from being cluttered with irrelevant information.
- The kernel has removed the use of the strncpy() function, which copies a specified number of bytes from the input string. The use of strncpy() posed risks of errors due to skipping the null terminator at the end of the string or adding unnecessary null padding. Instead of strncpy(), it is recommended to use the strscpy() and strscpy_pad() functions for copying null-terminated strings, as well as strtomem_pad(), memcpy_and_pad(), and memcpy() for copying strings of known fixed size. The work to eliminate the use of strncpy() in the kernel began in 2020 and required the acceptance of 362 changes from 70 developers.
- In the hypervisor KVM Support for Intel MBEC (Mode-Based Execution Control) and AMD GMET (Guest-Mode Execution Trap) extensions has been added, allowing separate execution rights for the kernel and user space in guest systems to be handled in memory translation tables. Previously, Intel and AMD hardware virtualization extensions only allowed memory pages to be marked as executable with a single bit, with software-level separation of rights for the kernel and user space at the hypervisor level. The use of MBEC and GMET enables the elimination of permission checks on the hypervisor side and significantly reduces the resource-intensive overhead of control transfer from the guest system to the hypervisor (VMexit).
- Network subsystem
- The implementation of the TCP-AO (TCP Authentication Option, RFC 5925) extension has been transitioned to use the new cryptographic library libcrypto, simplifying the code and improving efficiency. TCP-AO allows for the verification of TCP headers using MAC codes (Message Authentication Code), employing more modern algorithms HMAC-SHA-1-96 and AES-128-CMAC-96 instead of the previously available TCP-MD5 option based on the outdated MD5 algorithm.
- The number of subflows supported for Multipath TCP (MPTCP) connections has been increased from 8 to 64.
- Work continues to reduce the use of the global lock rtnl_lock in the kernel's networking stack.
- The implementation of the AppleTalk protocol stack has been removed from the kernel, which was used in Apple computers since 1985 and was replaced by TCP/IP in the 1990s. Additionally, components of ATM data transmission technology not related to PPPoATM have been removed, as well as ARCnet network interfaces based on ISA and PCMCIA buses, PCMCIA Bluetooth adapters, Chelsea TLS accelerators, code for integrating TLS with sockmap, and support for 5/10 MHz frequency ranges in the cfg80211/mac80211 wireless stack. Due to unresolved locking issues and lack of maintenance, the specific implementation of TLS processing acceleration based on the TCP Offload Engine has been removed (the more common TLS offload implementation remains). Code for 32-bit x_tables compatibility on 64-bit systems has been disabled and is planned for removal.
- Support for GRO (Generic Receive Offload) and GSO (Generic Segmentation Offload) mechanisms has been added to the pppoe driver for hardware acceleration of packet reassembly and segmentation. Using GRO and GSO significantly increases throughput for incoming traffic; for example, on MediaTek MT7621 devices in configurations with network address translation, maximum throughput increased from 130 Mbit/s to 630 Mbit/s.
- Hardware
- Initial support for HDMI 2.1 FRL (Fixed Rate Link) technology has been introduced in the AMDGPU driver, allowing for the transmission of uncompressed video at 4K/120Hz and 8K/60Hz quality. Previously, HDMI 2.1 support had not been implemented in open-source drivers for a long time due to compliance requirements from the HDMI Forum, but AMD has now managed to negotiate such an implementation.
- Support for configuring the background color by the display controller has been added to the i915 driver. The parameter pin_params.needs_low_address has been implemented.
- Work continues on the drm driver (Direct Rendering Manager) Xe for GPUs based on the Intel Xe architecture, which is used in Intel Arc graphics cards and integrated graphics starting from Tiger Lake processors. Initial support has been added for the Crescent Island (CRI) platform. A system controller has been implemented for dGPU platforms Xe3p.
- Issues with the NVIDIA GA100 GPU have been resolved in the Nouveau driver.
- The v3d driver has been enhanced with the capability to manage the power consumption of the v3D GPU on Raspberry Pi boards.
- Integration of the Nova driver components for NVIDIA GPUs equipped with GSP firmware has continued, starting with the NVIDIA GeForce RTX 2000 series based on the Turing microarchitecture. The driver is written in Rust. Support has been added for the NVIDIA GA100 GPU and the Hopper and Blackwell series.
- Support for ARM platforms, SoCs, and devices has been added: Apple t8122 (M3), Motorola Edge 30, Nothing Phone 3a, Google Pixel 3a XL, Qualcomm Dragonwing IPQ9650, Huawei Hawi, ZTE zx297520v3, Renesas R-Car M3Le, ASPEED AST27xx, Cortina Gemini, NXP i.MX6/8/9, NXP LX2160A, TI K3 AM62x.
At the same time, the Latin American Free Software Foundation has released a completely free kernel variant 7.2 — Linux-libre 7.2-gnu, cleaned from firmware elements and drivers containing non-free components or code segments with limited use defined by the manufacturer. Version 7.2 has removed blobs from the new drivers rt722-sdca and tac5xx2. The cleaning code in the drivers amdgpu, nova core, qcom iris, q6v5, iwlwifi, amdxnda, adreno, r8152, and mt792x has been updated. Adjustments have been made to the interfaces for loading firmware. The names of blobs in dts files (device tree) for ARM chips have been cleaned.
Source: opennet.ru
