Two experiments have highlighted AI’s expanding ability to interpret and mimic the human brain's process of audio and visual perceptions.
In a groundbreaking experiment, a team of scientists at the University of California, Berkeley, has recreated a snippet of a Pink Floyd song using the brain activity of people listening to it.
The 15-second audio piece, resembling 'Another Brick in the Wall, Part 1', is a pioneering demonstration of the intersection between neuroscience, artificial intelligence and music.
The research involved the brain scans of 29 patients undergoing epilepsy treatment who had electrodes implanted to locate the epicentres of drug-resistant seizures. Their brain activity was recorded while they listened to the Pink Floyd track.
How the brain listens to music
Explaining the complexity behind music listening, the brain essentially translates vibrations in the air into electrical signals, which neurons in the brain process to decode each lyric, melody and rhythm. The electrodes provided a unique view into this intricate process, allowing researchers to understand the relationship between sound and brain activity.
UC-Berkeley neuroscientist and an author of the study, Robert Knight, said, “Music adds the more emotional and melodic elements of speech. Understanding its processing in the brain is the first brick in the wall.”
He further compared the network of electrodes to piano keys with each electrode indicating how the brain responded to specific musical notes or rhythms.
Data from the study, collected between 2009 and 2015, sat untouched for almost a decade until postdoctoral researcher and musician Ludovic Bellier decided to decode it.
Bellier found that specific parts of the patients’ temporal lobes activated when the 16th notes of the guitar rhythm played. This previously undiscovered observation led researchers to the superior temporal gyrus, a brain region vital for processing sounds. Bellier pointed out that this region, especially on the right side, is seemingly responsible for perceiving rhythm.
Bellier transformed the brain activity data into music using an advanced machine-learning model, which converted the patterns into a spectrogram and, subsequently, an audio file. This audio was found to closely mirror the original Pink Floyd song.
This research may have significant implications for the medical field. Knight and Bellier believe their findings can revolutionise medical treatments, especially for individuals who have lost their ability to communicate.
Traditional brain signal-to-vocal translation machines tend to produce a robotic sound, but understanding the neural processing of music could add expressiveness to these devices.
The study hints at the potential for composing music via thought and could pave the way for inventions like 'a keyboard for the mind' to assist in translating patients' thoughts into words. Traditional brain signal-to-vocal translation machines tend to produce a robotic sound but understanding the neural processing of music could add expressiveness to these devices.
AI deciphers brain scans to reproduce sight
Another ground-breaking study brings us one step closer to understanding how the human brain translates visual stimuli into mental images.
The study by a group in Japan, which will be presented at an upcoming computer vision conference, reveals that AI can recreate images people have viewed using their brain scans.
By harnessing the power of a particular AI algorithm, Stable Diffusion, to decipher brain scans and reproduce images closely resembling what the participants saw.
Though the technique of using AI to interpret and reproduce brain scan-induced images is not new, with previous experiments focusing on human faces and landscape photos, what sets this research apart is the use of the Stable Diffusion algorithm.
Stable Diffusion
Developed by a German group in 2022, Stable Diffusion creates images from text prompts after digesting vast image-text datasets.
For this experiment, researchers fine-tuned the Stable Diffusion algorithm. By linking text descriptions of thousands of photographs to brain patterns recorded when participants viewed those images, the system showed a heightened ability to "decipher the brain".
Detailing the procedure, Yu Takagi, a systems neuroscientist at Osaka University involved in the study, explained that the AI algorithm used data from different brain regions, notably the occipital and temporal lobes, both pivotal in image perception.
The brain scans, sourced from functional magnetic resonance imaging (fMRI), tracked blood flow changes in active brain regions. These scans helped the AI interpret which parts of the brain process image contents (like objects or scenery) and which parts process layout aspects (like scale and position). Subsequently, the AI recreated images from these patterns.
Takagi shared that the AI-generated images initially resemble TV static. However, as the Stable Diffusion algorithm analysed a person's brain activity from viewing a photo and compared it to its training patterns, the noise transformed into identifiable features. Impressively, the new method required fewer resources, minimal fine-tuning and a smaller dataset compared to previous techniques.
However, it wasn't entirely without challenges. While the AI could capture the layout and perspective based on activity in the occipital lobe, it found it difficult to accurately reproduce specific objects from the real photos. For instance, instead of a clock tower, the AI might produce a vague shape.
To overcome this hurdle, the research team ingeniously incorporated keywords from image captions from an online dataset, which consisted of brain scans from four participants who viewed a set of 10,000 photos.
If a brain activity pattern linked to viewing a clock tower was identified, this keyword was fed into Stable Diffusion's text-to-image generator. Consequently, the AI produced an image with a clock tower, following the layout and perspective indicated by the brain pattern. This approach, Takagi and his team believe, allowed for the creation of images that closely mirrored the real photos.
These innovative studies signal AI's growing prowess in closely mirroring human cognition and perception and may have significant implications for the medical field. As AI evolves, its potential applications are boundless, with potential to revolutionise medical treatments.