Oracle has released the Unbreakable Enterprise Kernel 8

Oracle has released the first stable version of the Unbreakable Enterprise Kernel 8 (UEK R8), a variant of the Linux kernel developed for use in the Oracle Linux distribution as an alternative to the standard kernel package from Red Hat Enterprise Linux. The kernel is available only for x86_64 and ARM64 (aarch64) architectures. The source code of the kernel, including its breakdown into separate patches, is published in Oracle's public Git repository.

The Unbreakable Enterprise Kernel 8 package is based on Linux kernel 6.12 (the UEK R7 release was based on kernel 5.4, and the beta version of RHEL 10 offered kernel 6.11), which has been enhanced with new features, optimizations, and fixes, and has been tested for compatibility with most applications running on RHEL. It is optimized for use with Oracle industrial software and hardware. Installation and src packages for the UEK R8 kernel have been prepared for Oracle Linux 9.5 (there are no impediments to using this kernel in similar versions of RHEL, CentOS, Alma Linux, and Rocky Linux).

Key innovations of the Unbreakable Enterprise Kernel 8:

  • The separation of UEK kernel components into individual packages has been modified. Kernel modules are separated from the base kernel image and moved to collections provided in separate packages: kernel-uek-modules-core (base minimum), kernel-uek-modules (for servers), kernel-uek-modules-desktop, kernel-uek-modules-extra-netfilter, kernel-uek-modules-usb, and kernel-uek-modules-wireless. Related utilities have been moved from the base package kernel-uek-core to a separate package kernel-uek-tools. Configuration files listing modules prohibited from loading have been renamed from 'blacklist' to 'denylist' as part of an initiative to use inclusive terminology.
  • For ARM Ampere systems used in Oracle Cloud, a separate kernel build, kernel-uek64k, has been created, which increases the base memory page size from 4 KB to 64 KB.
  • Support has been added for hardware implementation of the EDMM (Enclave Dynamic Memory Management) mechanism provided in Intel SGX2 (Software Guard Extensions), allowing management of access to individual memory pages of the enclave and dynamically adding/removing memory pages for the enclave.
  • The Intel QAT driver supporting Intel Quick Assist Technology (QAT) devices has been updated to support the 4th generation Intel Xeon processors.
  • A split-lock detection system has been added, which occurs when accessing unaligned data in memory due to atomic instructions causing data to cross two CPU cache lines. Such locks lead to significant performance drops (up to 1000 cycles slower than atomic operations with data fitting into a single cache line).
  • A new method of protection against the Retbleed vulnerability in Intel and AMD CPUs has been implemented, utilizing call depth tracking, which does not slow down performance as much as the previous Retbleed protection.
  • On x86 systems, simultaneous activation of the CPU's secondary cores has been enabled, reducing the kernel boot time on systems with a large number of cores.
  • A new kernel command-line parameter 'ia32_emulation' has been added, allowing for enabling and disabling support for 32-bit mode emulation in kernels compiled for x86-64 architecture during the boot stage.
  • By default, the EEVDF (Earliest Eligible Virtual Deadline First) scheduler is now used instead of CFS (Completely Fair Scheduler). When selecting the next process to schedule, the new scheduler takes into account processes that have received fewer CPU resources or have been allocated unjustly more CPU time. In the first case, control is forcefully passed to the process, while in the latter, it is deferred. The old CFS scheduler used heuristics and fine-tuning to determine processes needing special attention, while the new scheduler tracks them more explicitly and does not require fine-tuning. EEVDF will reduce latency for tasks that CFS struggled to schedule.
  • The delivery of DTrace 2.0 dynamic debugging system has continued, which has been transitioned to use the eBPF kernel subsystem. DTrace 2.0 operates on top of eBPF, similar to how existing tracing tools in Linux operate with eBPF.
  • Up to 4096 virtual CPUs (VCPUs) are now allowed in the KVM hypervisor.
  • The use of KTLS, a kernel-level TLS protocol implementation, has continued.
  • The implementation of the RDRAND pseudo-random number generator, responsible for the operation of the device /dev/random, has been updated to use the BLAKE2s hash function instead of SHA1 for entropy mixing operations. This change has improved the security of the pseudo-random number generator. To speed up random number generation via the getrandom() system call, the vDSO (virtual dynamic shared object) mechanism has been employed, moving the system call handler from the kernel to user space to avoid context switches.
  • Support for the BIG TCP extension has been added, allowing an increase in the maximum TCP packet size to 4GB to optimize the performance of high-speed internal data center networks. This packet size increase, with a 16-bit field size in the header, is achieved through the implementation of 'jumbo' packets whose size in the IP header is set to 0, while the actual size is transmitted in a separate 32-bit field in an additional attached header.
  • An option SO_RESERVE_MEM has been implemented for network sockets, allowing a specific amount of memory to be reserved for the socket that will always remain available and not be reclaimed. Using this option increases performance by reducing memory allocation and return operations in the network stack, especially when faced with memory shortfalls in the system.
  • Performance optimization of the fq (Fair Queuing) packet scheduler has been conducted, resulting in a 5% increase in throughput under heavy loads in the tcp_rr (TCP Request/Response) test and a 13% increase in unlimited UDP packet streams.
  • Reorganization of kernel network structures has been carried out to improve CPU data caching efficiency, boosting the TCP stack performance on systems handling a large number of concurrent requests.
  • Support for ASMLib 3 library has been added for automated storage management in Oracle DBMS.
  • Work has been done to optimize performance and enhance the security of the io_uring asynchronous I/O mechanism. io_uring-based optimizations have been added for XFS and Ext4 file systems, enabling parallel direct file writes across multiple threads.
  • Improved support for the Btrfs file system. For devices that support trim/discard, the 'discard=async' mount option is enabled by default, allowing these operations to be performed immediately for all filesystems in asynchronous mode. Support for sending and receiving compressed data without conversions has been added. Support for writing blocks larger than 64 KB has been included. Quota accounting has been simplified. Support for mounting cloned devices has been implemented. Write checks in NOCOW mode have been improved (bandwidth increased by 9%). Mount options 'ignoremetacsums' and 'ignoresuperflags' have been added to ignore invalid metadata checksums and superblock flags. Task execution for removing devices, balancing, and redistributing blocks in parallel mode has been ensured.
  • In the XFS filesystem, the use of block sizes exceeding the page size is permitted. Large extent counters for very large virtual disks have been added. An atomic commit mode for file content has been introduced. An experimental fsck and online filesystem recovery implementation has been suggested.
  • NFS now defaults to using the READ_PLUS operation defined in the NFS 4.2 specification, which is applied for more efficient data reading from files containing holes.
  • The memory management system has been transitioned to using the folio data structure (folios of pages). Folios resemble compound pages but feature improved semantics and a clearer organization of operations.
  • The 'maple tree' data structure has been employed for memory mapping operations, positioning it as a more efficient replacement for the 'red-black tree'. The maple tree is a variant of the B-tree that supports range-based indexing and is designed for efficient cache usage. of modern processors.
  • In mmap, locks at the level of individual VMAs (Virtual Memory Areas) have been implemented, allowing for enhanced performance of multithreaded applications.
  • A ptdesc data structure has been added, optimizing operations with page tables by separating metadata structures from data for memory pages.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster