Recommended listening: Knife Party, Bonfire

In the previous instalment of this music technology series, we discussed the development of magnetic tape and how its use evolved from merely a medium for recording full performances to a way to play back shorter audio samples within a wider piece of music (often referred to as “samplers”). 

A major drawback of using tape in this way was the accumulation of background noise – once a sound is captured on tape, variations and imperfections in the magnetic particles inevitably lead to a small amount of background noise (“tape hiss”) when the sound is played back. The level of the noise could be reduced to an extent though improvements to the tapes themselves and the recording process, but it could not be eliminated completely.

The solution to this problem was digital audio encoding – i.e. taking a continuous “analog” signal (a waveform) and representing it as discrete data (a binary file). The innovation that enabled this is Pulse Code Modulation (“PCM”) and, as will be discussed later, its impact on the musical world was monumental. 

The first person to consider encoding audio as discrete data was Alec Reeves, who filed patents for the process from 1938 (e.g. see US2272070A – granted in 1942). According to the patent (see Figure 1), instead of transmitting the analog waveform itself (the black line), the amplitude of the original waveform was sampled (see circles on the black line) at predetermined time intervals (defined by a “sample rate”) so that “a signal code representing the nearest predetermined amplitude value above or below said instantaneous amplitude value” could be used to transmit the data. Crucially, by storing the underlying discrete data representing the signal rather than the signal itself, the data could be recovered without being as affected by noise, and the original signal could be recovered simply by “joining the dots”.

US2272070A, showing how a continuous audio wave (the black line) can be sampled at discrete locations (the circles). Amplitude data at each sample location can be transmitted, such that a receiver can reproduce the audio wave by joining the circles together.

However, while Reeves’ method was sufficient for its intended purpose of transmitting and reproducing intelligible voice messages, it lacked mathematical rigour and raised several questions. For example, was there more than one way to join the dots? If so, what is the best way to do so? How much data is lost in the process? How can this loss be minimised?

These questions (and many more) were answered by landmark patent filings from John Pierce, Claude Shannon, and Oliver Bernard (see US2437707 filed in 1945 and US2801281 filed in 1952). In these patents (and some papers published around that time), they proved that there was a unique way to connect the dots that preserved all frequency information below half the sample rate. Since human hearing only goes up to 20kHz (if you’re lucky), a sample rate of at least 40kHz is in principle completely sufficient to reliably record digital audio, though this comes with a couple of caveats. 

Firstly, the unique solution only matches the original signal if all frequencies above half the sample rate (known as the “Nyquist frequency”) are filtered out first. In US2437707, Pierce stated that if “it is desired to transmit all components up to 3000 cycles then there should be at least two samples per cycle for this highest frequency component”.  Failure to do so leads to “aliasing” where the original high-frequency information is erroneously reconstructed with funky-sounding metallic artifacts (more on this later!).

Secondly, the rounding of each sample to the nearest predetermined amplitude value does still introduce some background noise (“quantization error”).  However, the volume of this noise is very predictable based on the number of possible amplitude values (defined by the bit depth). Common formats nowadays are 16-bit (with 65536 unique amplitude values and 100 decibels of dynamic range) and 24-bit (with 16 million unique amplitude values and 144 decibels of available dynamic range). Especially when used with other processing like dithering (which counterintuitively involves adding noise), PCM encoding is capable of storing sound at quality levels far beyond what could be achieved on even the best tape recordings.

So, if PCM was figured out in the 1950s, why did it take until the 1990s for digital file formats to overtake analog ones? Well, the bottleneck was the digital hardware itself – using PCM in real-time required processors to access and manipulate thousands of possible amplitude values thousands of times per second. As a result, it took a few decades for the music industry to discover the full potential of digital audio, with several uses and applications being revealed as the speeds of processors and the costs of digital storage gradually improved.

A recording medium

The most obvious initial use for PCM in the world of music was as a way to replace tape as a medium for recording, storing and distributing music. Several classical works were recorded this way in the early 1970s, with the first digital releases of pop music coming in 1979 with “Bop Till You Drop” by Ry Cooder, shortly followed by Stevie Wonder’s “Journey Through the Secret Life of Plants”. 

CDs (which use 44.1kHz, 16-bit PCM) surged in popularity as a distribution medium through the 80s and 90s, and nowadays all audio files that are downloaded and streamed use PCM in some form (e.g. mp3, wav, aiff, or flac files) with the differences largely being the type of data compression that is subsequently applied. 

Samplers

Much like Mellotron mentioned in the previous article, engineers soon figured out how to store and trigger digital samples so that they could be played like musical instruments. The first commercially available digital sampler (using rudimentary 12-bit 22kHz PCM) was the Computer Music Melodian in 1976, and this was used by Stevie Wonder on the album “Journey Through the Secret Life of Plants”, mentioned above. 

Shortly afterwards, in 1979, the iconic Fairlight CMI was released. This instrument was surprisingly advanced for the time and allowed sounds to be recorded and edited via a touch screen. Many musicians loved the Fairlight (e.g. Kate Bush used it extensively, including for the first sound you hear in “Running Up That Hill” in 1985), though it still had critics (e.g. Phil Collins, who put a notice on the sleeve of his 1985 album “No Jacket Required” that stated that “there is no Fairlight on this record”). 

Digital sampling was adopted in various drum machines, such as the Linn LM-1 in 1980 (used by Prince in “When Doves Cry”), the Roland TR-909 in 1983 (as used by Madonna in “Vogue”) and the E-mu SP-1200 in 1989 (used throughout the Hip-Hop genre, such as in Gang Starr’s “Take it Personal”).

Granular and wavetable synthesis

The applications above were already possible to some extent using existing tape-based methods. However, PCM really came into its own when processors became sufficiently advanced to allow them to precisely manipulate waveforms at a more granular level. 

For example, rather than playing back the full recording of a sample, small fragments of the audio (e.g. from 1 to 100ms) could be repeated, overlapped and otherwise manipulated in an approach known as “granular synthesis”. This was first implemented in real-time by Barry Traux in 1986, and while not used extensively in popular music, it was adopted by more experimental musicians to produce unusual morphing soundscapes

Taking this idea even further, individual cycles of a PCM waveform could be stored in a lookup table, such that the active waveform could be selected and changed in real time. This type of synthesis is known as “wavetable synthesis”, and enables interesting and unique sounds to be produced without needing vast amounts of data storage while still enabling real-time control. 

Throughout the 2000s and 2010s, software-based wavetable synthesisers became particularly common, with the most iconic arguably being Native Instruments’ Massive released in 2006. Massive was so popular that some of its individual wavetables became well-known in their own right. For example, the ‘modern talking’ wavetable on Native Instruments Massive was (over)used so much that it became infamous among music producers

Demonstration of NI Massive’s “modern talking” wavetable, and how movement of the wavetable position dial (in the top left) can lead to a bass sound just like the one in Nicky Romero’s 2011 track Toulouse.

Several songs have already been mentioned that demonstrate the various uses of PCM. More broadly, anything released in recent decades is practically guaranteed to have used digital processing at least once during the production process. So why suggest Knife Party’s 2013 dubstep tune, “Bonfire”?

Well, like the music mentioned above, it was distributed as a digital file format. Likewise, it also made heavy use of sampling, such as for elements like the drums and the vocals, and these would have been arranged on a computer using a Digital Audio Workstation (DAW). Furthermore, Knife Party were known users of the wavetable synthesiser NI Massive, and it is likely that this would have been used for several of the sounds in Bonfire.

However, this song differs from all the above mentions in that it deliberately (ab)uses the PCM encoding in an audible way to provide a signature part of some of the bass sounds. More specifically, while the “aliasing” mentioned earlier was something to be avoided during production of most music genres, the resulting “funky sounding metallic artifacts” were considered a definite positive by electronic music producers of the 2010s. Below is a recreation of the sound you hear 1 minute into the song. When the sample rate is artificially reduced to 1.59 kHz without filtering out any frequencies above half the sample rate, a characterful “yah” sound is created. 

Demonstration of how sample rate reduction can transform a fairly unexciting filtered sawtooth wave into a crunchy robotic bass sound, just like the one that occurs 1 minute into Bonfire.

Although sample rate reduction was not completely unique to electronic music (e.g. see “We’re In This Together” by Nine Inch Nails in 1999), it was certainly the genre where its use was most ubiquitous.

It is hoped that the discussion above gives you a taste of the ubiquity of PCM within modern music, and more broadly that this series of articles has demonstrated the transformative role that technology has played in the creation of music. A hundred years ago, high-fidelity music could only be experienced in a live setting. Now, streaming services like Spotify have upward of 100 million songs available at the touch of a button, with this number growing by tens of thousands every day. These songs include not just those on big-budget albums from well-known artists, but also those created by amateur musicians (myself included!) and make use of sounds, musical instruments and production techniques that would have been impossible only a few decades before.

If you would like to discuss anything in any of the articles further, or you have an invention that you would like to protect, then please contact the author, or get in touch with our patents team at gje@gje.com.