NVIDIA Releases Source Code for StyleGAN3, a Machine Learning System for Face Synthesis

NVIDIA has released the source code for StyleGAN3, a machine learning system based on a generative adversarial network (GAN) aimed at synthesizing realistic images of human faces. The code is written in Python using the PyTorch framework and is distributed under the NVIDIA Source Code License, which imposes restrictions on commercial use.

Pre-trained models are also available for download, trained on the Flickr-Faces-HQ (FFHQ) collection, which includes 70,000 high-quality (1024×1024) PNG images of human faces. Additionally, models based on collections like AFHQv2 (animal face photographs) and Metfaces (images of human faces from classical portraits) are available. While the focus is on faces during development, the system can also be trained to generate any objects, such as landscapes and cars. Tools for independently training the neural network with your own image collections are also provided. One or multiple NVIDIA graphics cards are required (Tesla V100 or A100 GPU recommended), with at least 12 GB of RAM, PyTorch 1.9, and CUDA 11.1+ toolkit. A special detector is being developed to identify the artificial nature of the generated faces.

The system allows for the synthesis of a new face image through the interpolation of features from multiple faces, combining their characteristic traits while adapting the final image to a desired age, gender, hair length, smile type, nose shape, skin color, glasses, and photo angle. The generator treats the image as a collection of styles, automatically separating distinctive details (freckles, hair, glasses) from common high-level attributes (pose, gender, aging changes) and allows for arbitrary combinations with dominant properties identified through weighting coefficients. As a result, images are generated that are indistinguishable from real photographs.

NVIDIA Releases Source Code for StyleGAN3, a Machine Learning System for Face Synthesis

The first version of the StyleGAN technology was published in 2019, followed by an improved edition, StyleGAN2, in 2020, which enhanced image quality and eliminated some artifacts. However, the system remained static, meaning it did not allow for realistic animation and facial movement. The main goal in developing StyleGAN3 was to adapt the technology for its application in animation and video.

StyleGAN3 features a reworked image generation architecture that eliminates aliasing and introduces new training scenarios for the neural network. It includes new utilities for interactive visualization (visualizer.py), analysis (avg_spectra.py), and video generation (gen_video.py). The implementation also reduces memory consumption and accelerates the training process.

NVIDIA Releases Source Code for StyleGAN3, a Machine Learning System for Face Synthesis

A key feature of the StyleGAN3 architecture is the transition to interpreting all signals in the neural network as continuous processes, allowing manipulation of relative positions when forming details, which are not tied to absolute pixel coordinates in the image but are anchored to the surface of the depicted objects. In StyleGAN and StyleGAN2, the pixel binding during generation led to problems in dynamic visualization; for instance, when the image moved, there was a misalignment of small details, such as wrinkles and hairs, which moved separately from the rest of the face. In StyleGAN3, these issues are resolved, making the technology suitable for video generation.

Additionally, it is worth noting the announcement of the creation by NVIDIA and Microsoft of the largest language model, MT-NLG, based on a deep neural network with a transformer architecture. The model encompasses 530 billion parameters, and a cluster consisting of 4480 GPUs was used for training (560 DGX A100 with 8 A100 80GB GPUs each). servers The model is intended for applications in natural language processing tasks, such as predicting the completion of unfinished sentences, answering questions, reading comprehension, generating conclusions in natural language, and disambiguating word meanings.

NVIDIA Releases Source Code for StyleGAN3, a Machine Learning System for Face Synthesis


Source: opennet.ru
Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster