The Stable Diffusion machine learning system has been adapted for music synthesis.

The Riffusion project develops a variant of the machine learning system Stable Diffusion, adapted for generating music instead of images. Music can be synthesized from a natural language textual description or based on a provided template. The components for music synthesis are written in Python using the PyTorch framework and are available under the MIT license. The interface wrapper is implemented in TypeScript and is also distributed under the MIT license. The trained models are open under the permissive Creative ML OpenRAIL-M license, allowing for commercial use.

The project is interesting as it continues to use 'text-to-image' and 'image-to-image' models for generating music, but manipulates spectrograms as images. In other words, the classic Stable Diffusion is trained not on photographs and pictures, but on images of spectrograms, reflecting the change of frequency and amplitude of sound waves over time. Consequently, a spectrogram is also generated as output, which is then transformed into a sound representation.

The Stable Diffusion machine learning system has been adapted for music synthesis.

The method can also be used to modify existing sound compositions and synthesize music by template, similar to how images are modified in Stable Diffusion. For example, during generation, spectrogram samples with a reference style can be set, different styles can be combined, smooth transitions from one style to another can be executed, or changes can be made to existing sounds to address tasks such as increasing the volume of individual instruments, changing the rhythm, and replacing instruments. Samples are also used to generate long-playing compositions, composed of a series of segments that are close to each other and slightly change over time. Separately generated segments are combined into a continuous stream using interpolation of the model's internal parameters.

The Stable Diffusion machine learning system has been adapted for music synthesis.

The Fourier window transform is used to create a spectrogram from sound. When reconstructing sound from the spectrogram, a phase determination problem arises (the spectrogram contains only frequency and amplitude); to reconstruct it, the Griffin-Lim approximation algorithm is used.



Source: opennet.ru
Buy reliable website hosting with DDoS protection, VPS VDS servers 🔥 Buy reliable website hosting with DDoS protection, VPS VDS servers | ProHoster