Release of Kata Containers 3.4 with Virtualization-Based Isolation

The release of Kata Containers 3.4 has been published, advancing the stack for container execution using isolation based on full virtualization mechanisms. The project was created by Intel and Hyper by merging the technologies of Clear Containers and runV. The project's code is written in Go and Rust and is distributed under the Apache 2.0 license. The development of the project is overseen by a working group established under the auspices of the independent OpenInfra Foundation, which includes companies such as Canonical, China Mobile, Dell/EMC, EasyStack, Google, Huawei, NetApp, Red Hat, SUSE, and ZTE.

At the core of Kata is a runtime that allows the creation of compact virtual machines running with a fully-fledged hypervisor instead of traditional containers that use a shared Linux kernel and are isolated using namespaces and cgroups. The application of virtual machines provides a higher level of security, protecting against attacks resulting from the exploitation of vulnerabilities in the Linux kernel.

Kata Containers is geared towards integration into existing container isolation infrastructures with the possibility of using such virtual machines to enhance the security of traditional containers. The project provides mechanisms to ensure compatibility of lightweight virtual machines with various container isolation infrastructures, container orchestration platforms, and specifications such as OCI (Open Container Initiative), CRI (Container Runtime Interface), and CNI (Container Networking Interface). Tools for integration with Docker, Kubernetes, QEMU, and OpenStack are available.

Integration with container management systems is achieved through a layer that simulates container management, which accesses the controlling agent in the virtual machine through a gRPC interface and a special proxy. Inside the virtual environment, which is launched by the hypervisor, a specially optimized Linux kernel is used, containing only the minimum set of necessary features.

Dragonball Sandbox (a KVM-based hypervisor optimized for containers) is supported as the hypervisor, along with QEMU tools, Firecracker, and Cloud Hypervisor. The environment includes an initialization daemon and an agent. The agent facilitates the execution of user-defined container images in OCI format for Docker and CRI for Kubernetes. When used with Docker, a separate environment is created for each container. the virtual machine, meaning that the environment running on top of the hypervisor is used for nested container execution.

Release of Kata Containers 3.4 with Virtualization-Based Isolation

To reduce memory usage, the DAX mechanism (direct access to the file system bypassing the page cache without using block device levels) is applied, while the KSM (Kernel Samepage Merging) technology is used for deduplicating identical memory regions, allowing for resource sharing of the host system and connecting a shared template of the environment to different guest systems.

To organize access to container images, the Nydus file system is utilized, which uses content-addressing for efficient collaboration with standard images. Nydus supports on-the-fly image loading (only loading when needed), provides deduplication of duplicate data, and can use various backends for actual storage. Compatibility with POSIX is available (similarly to Composefs, the Nydus implementation combines the capabilities of OverlayFS with EROFS or a FUSE module).

In the new version:

  • The Dragonball virtual machine manager has added support for hot-plugging of GPUs and the ability to use Memory-Type Range Registers (MTRR) for accessing physical memory areas.
  • In runtime-rs, the Rust language runtime implementation, full processing of threads, pid, and tid has been ensured, and the qemu driver, used in systems with s390 architecture (IBM Z), has been redesigned.
  • The service for creating snapshots using the Nydus file system has been updated.
  • In the container image management service, memory operation efficiency has been improved.
  • By default, mounting of the cgroups-v2 hierarchy during boot has been enabled using systemd.
  • The ability to define a timeout to limit the time for retrieving very large images in guest systems has been added.
  • Support for building the OPA agent (Open Policy Agent) for ppc64le and s390x architectures has been added.

Source: opennet.ru

Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster