KIAS
Transdisciplinary
Research Program
KIAS
Research Background and Motivation
Most people frequently encounter music in their daily lives and feel its emotional power as a personal experience. In modern society, while music is often consumed as a commodity of popular art, it frequently evokes emotions and memories by being strongly associated with specific events or experiences undergone by an individual. On a social level, the music of a generation can have a profound influence on the formation of that generation's identity and the way emotions are communicated within society.
IProfessor Daniel Levitin, a neuroscientist and accomplished musician, in his seminal work This Is Your Brain on Music (Plume/Penguin, 2007), elucidates the reasons why humans are drawn to music and explores its profound influence, weaving together his personal musical experiences with the analytical language of science. Drawing on everyday phenomena like musical preferences, as well as findings from cognitive psychology experiments, neuroanatomy, and neurochemistry, he attempts to provide an integrated explanation of how our brains process musical elements—such as interval, pitch, rhythm, and timbre—and the mechanisms behind the emotional and cognitive responses evoked by music. In short, Levitin's book offers remarkable insights into the interaction between music and humans. Although not detailed here, the way the brain perceives basic musical elements like pitch, interval, rhythm, timbre, and harmony is remarkably aligned with the ideas of modern artificial intelligence learning theories.
The first step in the interaction between music and humans is the process in which sound wave signals entering our auditory organs—the ears—are physically converted, amplified, and detected. Regarding this mode of auditory sensing, Dennis Gabor, a Nobel laureate in physics and an amateur musician, introduced the concept of "acoustical quanta" and proposed an auditory theory resembling Heisenberg's uncertainty principle. Gabor's phonemes are mathematically represented in the form of Gaussian sinusoidal waves and physically represent pulse signals in time-space or signals occupying specific bands in frequency-space. This simple concept became the foundation of modern electronic acoustics theory and, in particular, the basis for the electronic music generation method known as Granular Synthesis. In contrast to his significant contributions to modern electronic music, there has not yet been detailed research on how Gabor's auditory theory connects to the structure and operation of the inner ear as revealed by modern science.

Figure 1 Mathematical representation of a quantum mechanical acoustic element (Nature 1947). The minimum of sound perception proposed by Gabor
It can be quantified as a unit by the period, time width, maximum amplitude, and phase of a Gaussian sine wave.
Meanwhile, Dennis Gabor also took an interest in vision and devised a visual training method called the Gabor patch. While the phonemes introduced above are one-dimensional signals, the Gabor patch takes the form of a Gaussian sinusoidal pulse in two-dimensional space; the mode of each patch is defined by the periodic pulse width of the sinusoidal wave and the spatial direction of the pattern. Gabor's visual training utilizes patches of various modes to induce the use of eye muscles and stimulate the brain's visual information processing circuits. In particular, regarding the latter, the neuroscientific basis for such a training method was established by revealing that the receptive patterns of the brain's visual neurons are similar to the patterns of the Gabor patch. In Japan, it was recognized as an ophthalmic prescription for improving eyesight, and at one time, Gabor patch-based vision training books were published and became popular among the public. Recently, a mobile game app called GaborGame has been released.

Figure 2 Example of a Gabor pattern. During visual training, multiple patterns are given simultaneously, and the same pattern is found or
Training is conducted by distinguishing different patterns.
This interdisciplinary research group aims to devise auditory training methods similar to Gabor patches by utilizing Gabor’s concept of acoustics, thereby understanding the processes by which hearing functions and developing methodologies to stimulate auditory elements. Simultaneously, we intend to analyze whether elements connected to the physical operation of the auditory organs exist within existing classical music performances (particularly piano music). Through this, we aim to present a new perspective on exploring the relationships between music and memory (how music enables the recall of events and experiences), music and the body (the relationship between music and physical responses such as changes in heart rate and respiration), and music and mental states.
Research Scope and Content
The characteristics of human sound perception introduced by Gabor in his paper can be summarized in two points: namely, the limits of frequency resolution ( ∆f ) and parallax resolution ( ∆t ), expressed in a form resembling Heisenberg's uncertainty principle ( ∆f x ∆t ). >1), and the second stage is frequency discrimination ability (the second mechanism for searching the maximum frequency component). It is a reasonable hypothesis to assume that these characteristics are closely related to the physical operation of the auditory organs. Of course, understanding the entire auditory process, including cognition and memory in the brain, in detail is a very difficult and specialized field that falls outside the scope of research by this interdisciplinary research group.
Based on the above hypothesis, our research team intends to carry out activities in the form of literature review, modeling, and experimentation on the following topics.
(1) Exploration of the physical structure and characteristics of the auditory organs
The organ of Corti within the cochlea consists of approximately 3,500 cellular units, each composed of three outer hair cells (actuators), one inner hair cell (sensor), and a basilar membrane (resonator). It is known that when sound waves of the frequency assigned to each auditory unit are detected, the basilar membrane resonates; the outer hair cells near the corresponding membrane act as actuators or attenuators to control this resonance and maintain the vibration for a certain period. The vibration of the basilar membrane is detected by the inner hair cells acting as sensors and transmitted to the brain via the auditory nerve. While Helmholtz's resonator is similar to the organ of Corti in that it enables real-time frequency precision detection, it is much simpler in that its resonance conditions are fixed. This human auditory system can be viewed as a system in which real-time Fourier transform and adaptive signal amplification functions are physically implemented, and it is fundamentally different from a microphone that operates as a single sensor. This implies that the way multidimensional auditory signals transmitted to the brain are processed can be very different from the conventional way one-dimensional speech signals are processed, which are converted into electrical signals through a microphone.
This interdisciplinary research group aims to establish a fundamental understanding of the auditory system by conducting a literature review on the physical structure of each component constituting the auditory system, the range of physiological changes, and the modes of medical hearing impairment, as well as through consultation with medical experts. In particular, the group investigates whether the vibrational state of the basilar membrane can be classified as a physical "excited state" and develops a physical model for basilar membrane vibration.

Figure 3 Cross-sectional view of the organ of Corti inside the cochlea.
(2) Mathematical expression of the signal generated by the inner hair cell (sensor) population
The signals transmitted from the approximately 3,500 auditory units that make up the cochlea to the auditory nerve can be thought of as points in a multidimensional data space. We develop a method to mathematically analyze the patterns of acoustic data arrangement in this multidimensional space and to visually represent them.
(3) Model for recognizing harmonics, chords, timbres, and dissonances
When a sound containing harmonics enters the ear, it can be assumed that multiple auditory units within the organ of Corti become excited through the process described above. In the case of a sound without harmonics, only a single auditory unit will become excited. In the case of chords, it is presumed to be associated with the spatial pattern of the auditory units that become excited among the group of 3,500 auditory units. To distinguish between consonant and dissonant chords, we explore whether the mathematical relationship between two sounds can be defined as a "distance (measure)" or pattern in a multidimensional data space.
Timbre is also expected to be explained by the spatial patterns of excited auditory units and the temporal changes of those patterns in a multidimensional data space.
In the same manner, an analysis is performed in a multidimensional data space for imperfect intervals inevitably induced by harmonic series. The results of the analysis are compared with the definition of dissonant intervals in music theory. Furthermore, the relationship with cognitive dissonance (Auffassung dissonance: a phenomenon where an interval is physically consonant but perceived as dissonant due to a musical context: perfect fourths, 46th chords, etc.) and the relationship with previously known auditory, psychological, and chaos-theory analysis results are explored.
(4) Classical Music Audio Analysis – Chords vs. Arpeggios
It is known that the time required for an auditory unit to change from a ground state to an excited state is about 10 ms, and the vibration of the basilar membrane gradually increases until it reaches saturation by about 250 ms; however, our brain perceives the presence of sound earlier, within about 50-100 ms, and then judges the intensity of the sound over a period of about 200 ms. This point, along with the fact that the length of a 1/8 note is typically about 250 ms (of course, the numerical closeness between the two is likely coincidental and lacks scientific basis), allows us to consider an interesting and useful concept. That is, it is expected that the "clarity" of piano performance can be defined as an indicator capable of distinguishing articulation during performance. We intend to devise a mathematical expression method that can define clarity by quantifying data relationships in a multidimensional data space, such as being able to play a key for more than 100 ms without mixing with the next note, and in the case of chords, being able to play keys with very high simultaneity.
(5) Auditory stimulation system based on the concept of quantum mechanical acoustics (AQ)
We design and fabricate an acoustical quanta stimulation (AQS) system that stimulates auditory units within the cochlea by generating quantum mechanical acoustics. Factors to consider in designing the AQS system include (1) the physical operation of auditory cells, (2) fundamental limitations of time-frequency resolution inherent in the auditory system itself, (3) basic limitations/phenomena of hearing (critical band, masking, third sound: combinational sound), (3) psychophysical limitations of discrimination performance (just noticeable difference, JND), and (4) physiological indicators reflecting peripheral function (otoacoustic emissions).
Separately from the AQS system, we will develop an acoustical quanta (AQ)-based training app. The development of the smartphone app will be planned and executed as an undergraduate research project.
(6) Auditory perception experiment
This study measures acoustic-level auditory recognition ability using an AQS system on a population with differences in auditory perception and tests the effectiveness of AQS training. Even when the auditory system function, which contributes to sound sensation and cognition, is generally normal, differences such as absolute pitch, relative pitch, and tone deafness are understood to stem primarily from differences in the information processing and memory systems of the cerebral cortex; therefore, the purpose is to test whether AQS training brings about significant changes in the cerebral cortex.
Auditory perception abilities are measured in the same manner for professional classical musicians and compared to those of the general public. From a cognitive science perspective, musicians' "precise listening" may be more sensitive or precise than that of the general public, even in subcortical regions (particularly the auditory brainstem). This is reported as differences in indicators such as the Auditory Brainstem Response (ABR), which measures how quickly and consistently the brainstem responds to the onset of sound, and the Frequency-Following Response (FFR), which measures how precisely the brainstem/sub-auditory system tracks the periodicity and pitch structure of sound. Furthermore, numerous studies have reported that while professional musicians tend to have efficient long-term representations and long-term memory–working memory interaction, individuals with amusia may be relatively vulnerable in pitch discrimination, interval processing, or certain stages of memory/mapping. Based on this, we aim to develop FFR, pitch discrimination, and interval discrimination indices based on acoustics.