
Expressing what words cannot convey; feeling a myriad of emotions intertwining in a whirlwind of feelings; breaking free from the earth, the sky, and even the universe itself, embarking on a journey where there are no maps, no roads, no signs; inventing, narrating, and experiencing an entire story that will always remain unique and unmatched. All of this is made possible by music — an art form that has existed for many thousands of years, delighting our ears and hearts.
However, music, or more precisely musical compositions, can serve not only for aesthetic pleasure but also for transmitting encoded information intended for some device and unnoticed by the listener. Today, we will get acquainted with a rather unusual study in which graduate students from the Swiss Federal Institute of Technology Zurich were able to subtly embed certain data into musical works, making music itself a channel for data transmission. How exactly did they implement their technology, how do melodies with and without embedded data differ, and what did practical tests reveal? We will learn about this from the researchers' report. Let’s go.
The foundation of the research
The researchers refer to their technology as an acoustic data transmission technique. When a speaker plays a modified melody, a person perceives it as ordinary, while, for example, a smartphone can read the encoded information between the lines, or rather between the notes, if you will. The most important aspect of implementing this data transmission method that the scientists (the fact that these guys are still graduate students does not prevent them from being scientists) emphasize is the speed and reliability of transmission while maintaining these parameters regardless of the chosen audio file. Psychoacoustics, which studies the psychological and physiological aspects of how humans perceive sounds, helps tackle this challenge.
The core of acoustic data transmission can be defined as OFDM (Orthogonal Frequency Division Multiplexing), which, along with adaptation of subcarriers to the original music over time, allows for the maximum utilization of the transmitted frequency spectrum for information transfer. This has enabled a transmission speed of 412 bits/s over a distance of up to 24 meters (error rate < 10%). Practical experiments involving 40 volunteers confirmed that it is nearly impossible to discern the difference between the original melody and the one into which the information was embedded.
Where can such technology be practically applied? Researchers have their version of the answer: virtually all modern smartphones, laptops, and other pocket devices are equipped with microphones, and many public places (cafés, restaurants, shopping centers, etc.) have speakers playing background music. Data for connecting to Wi-Fi networks can be embedded in this background melody without requiring any additional actions.
We have clarified the common features of acoustic data transmission; now let’s move on to a detailed study of the structure of this system.
System Description
Data embedding into the melody occurs through frequency masking. During time intervals, masking frequencies are identified, and OFDM subcarriers close to these masking elements are filled with data.

Image 1: transformation of the original file into a composite signal (melody + data) transmitted through speakers.
To begin with, the original audio signal is divided into sequential segments for analysis. Each segment (Hi) of L = 8820 samples, equal to 200 ms, is multiplied by a window* to minimize boundary effects.
Window* is a weighting function used to control the effects caused by the presence of side lobes in spectral estimations.
Next, the dominant frequencies of the original signal were detected in the range from 500 Hz to 9.8 kHz, which allowed for the acquisition of masking frequencies fM,l for this segment. Additionally, data transmission was performed in a small range from 9.8 to 10 kHz to establish the location of the subcarriers in the receiver. The upper limit of the frequency range used was set at 10 kHz due to the low sensitivity of smartphone microphones to high frequencies.
Masking frequencies were determined individually for each analyzed segment. Using the HPS (Harmonic Product Spectrum) method, three dominant frequencies were established, which were then rounded to the nearest notes of the harmonic chromatic scale. Thus, the fundamental notes fF,i = 1…3, lying between the keys C0 (16.35 Hz) and B0 (30.87 Hz), were obtained. Given that the fundamental notes are too low for data transmission, their higher octaves 2kfF,i were calculated in the range of 500 Hz to 9.8 kHz. Many of these frequencies (fO,l1) were more pronounced due to the nature of HPS.

Image №2: calculated octaves fO,l1 for fundamental notes and harmonics fH,l2 of the strongest tone.
The set of octaves and harmonics was subsequently used as masking frequencies, based on which the OFDM subcarrier frequencies fSC,k were obtained. Two subcarriers were inserted below and above each masking frequency.
Next, the spectrum of the audio segment Hi was filtered at the frequencies of the subcarriers fSC,k. Based on the information bits in Bi, an OFDM symbol was created, enabling the composite segment Ci to be transmitted through the speaker. The magnitudes and phases of the subcarriers must be selected in such a way that the receiver can extract the transmitted data while the listener does not notice changes in the melody.

Image №3: segment of the spectrum and frequencies of the subcarriers of segment Hi of the original melody.
When the audio signal encoded with information is played through the speakers, the microphone of the receiving device records it. To locate the starting positions of the embedded OFDM symbols, the recordings must first be passed through bandpass filtering. This extracts the upper frequency range, where there are no musical interference signals between the subcarriers. The beginning of the OFDM symbols can be found using the cyclic prefix.
After detecting the start of the OFDM symbols, the receiver obtains information about the most dominant notes by decoding the upper frequency region. Moreover, OFDM is quite resistant to the influence of narrowband interference sources, as they affect only some of the subcarriers.
Practical tests
The sound source for the modified melodies was the KRK Rokit 8 speaker, and the receiving side was represented by the Nexus 5X smartphone.

Image No. 4: the difference between actual OFDM manifestations and correlation peaks measured indoors at a distance of 5 m between the speaker and the microphone.
Most OFDM points lie in the range from 0 to 25 ms, so a permissible start can be found within the cyclic prefix of 66.6 ms. Researchers note that the receiver (in this experiment, the smartphone) accounts for the fact that OFDM symbols are transmitted periodically, which improves their detection.
The first thing that needed to be tested was the effect of distance on the bit error rate (BER). Three tests were conducted in different types of rooms: a carpeted hallway, a room with linoleum flooring, and an auditorium with a wooden floor.

The song 'And The Cradle Will Rock' by Van Halen was chosen as the 'test subject'.
The volume was set so that the measured level at a distance of 2 m from the speaker was 63 dB according to the smartphone.

Image No. 5: BER indicators as a function of the distance between the speaker and the microphone (blue line — auditorium, green — hallway, orange — office).
In the hallway, sound at 40 dB was captured by the smartphone at a distance of up to 24 meters from the speaker. In the auditorium at a distance of 15 m, the sound was 55 dB, and in the office at a distance of 8 meters, the level of sound perceived by the smartphone reached 57 dB.
Due to the audience and office being more reverberant, the late echo signals of OFDM symbols exceed the length of the cyclic prefix and increase the BER.
Reverberation* — the gradual reduction in sound intensity due to its multiple reflections.
Next, the researchers demonstrated the versatility of their system by applying it to 6 different songs across three genres (see table below).

Table No. 1: Songs used in the tests.
Also, through the data in the table, we can see the bit rates and error coefficients for each song. The data transfer rates differ because differential BPSK (phase modulation) performs better when the same subcarriers are used. This is possible when adjacent segments contain identical masking elements. Continuously loud songs provide an optimal base for data hiding, as masking frequencies are more prominently present across a wide frequency range. Fast-changing music can only partially mask OFDM symbols due to the fixed window length for analysis.
Next, the system was tested by participants who had to determine which melody was original and which was modified with embedded information. For this, 12-second excerpts from the songs in Table No. 1 were placed on a special website.
In the first experiment (E1), each participant was given either a modified or original fragment to listen to, and they had to decide whether this fragment was original or altered. In the second experiment (E2), participants could listen to both versions as many times as they wanted and then determine which one was the original and which was modified.

Table No. 2: Results of experiments E1 and E2.
The results of the first experiment include two metrics: p(O|O) — the percentage of participants who correctly identified the original melody, and p(O|M) — the percentage of participants who identified the modified version of the melody as the original.
Interestingly, some participants reportedly considered certain modified melodies to be more original than the original itself. The average result from both experiments indicates that the average listener will not notice a difference between the original melody and the one that has data embedded in it.
Naturally, music connoisseurs and musicians will be able to catch some inaccuracies and suspicious elements in the modified melodies, but these elements are not significant enough to cause discomfort.
And now we can participate in the experiment ourselves. Below are two versions of the same melody — the original and the modified. Can you hear the difference?
vs
For a more detailed understanding of the nuances of the research, I recommend checking out the from the research group.
You can also download a ZIP archive of the audio files of the original and modified melodies used in the study at .
Epilogue
In this paper, PhD students from the Swiss Federal Institute of Technology Zurich described an astonishing data transmission system within music. They applied frequency masking, allowing data to be embedded in a melody played through a speaker. This melody is perceived by a device's microphone, which recognizes the hidden data and decodes it, while the average listener remains unaware of any difference. In the future, the team plans to develop their system, selecting more sophisticated methods for embedding data into audio.
Whenever someone invents something unusual, and importantly, functional, we always celebrate. But even more joy comes from the fact that this invention was created by young people. Science knows no age limits. If young people find science boring, it means it is not being presented from the right angle, so to speak. After all, as we know, science is a wonderful world that never ceases to amaze.
Friday off-topic:

Since we are talking about music, more specifically rock music, here’s a wonderful journey through the realms of rock.

Queen, “Radio Ga Ga” (1984).
Thank you for your attention, stay curious, and have a great weekend, everyone! 🙂
Thank you for staying with us. Do you enjoy our articles? Would you like to see more interesting materials? Support us by placing an order or recommending us to your friends. 30% discount for Habr users on a unique entry-level server designed by us for you: (options available with RAID1 and RAID10, up to 24 cores and up to 40GB DDR4).
Dell R730xd for half the price? Only with us in the Netherlands! Dell R420 — 2x E5-2430 2.2GHz 6C 128GB DDR3 2x960GB SSD 1Gbps 100TB — from $99! Read about how
Source: habr.com
