When playing back an audio sent by oneself, such as on WhatsApp, many people notice that the voice sounds different, often higher pitched, raising questions about how others perceive it. The reason for this disparity lies in the fact that the experience of hearing one's own spoken voice differs from the experience of hearing a recorded voice.
When we speak, our voice follows two distinct paths. One is sound propagation through the air, where sound waves reach the eardrums, which vibrate and send neural signals to the brain for auditory interpretation. The second path involves vibrations transmitted through the bones of the head. The skull acts as a resonance chamber, causing the inner ear, especially the cochlea, to vibrate directly.
This bone conduction modifies the apparent spectrum of the voice, resulting in a greater relative contribution of low frequencies, which gives the perceived timbre more deep and full characteristics. Thus, when we are speaking, we hear the sum of these two simultaneous transmissions: one aerial and one bone, forming the voice we consider ours.
However, when listening to a recording of one's own voice, we only have access to the sound transmission through the air. Since the information received by the ears does not correspond to the combination of the two factors present in natural speech, the perceived voice is different. Additionally, recording equipment is not ideal; it processes and compresses the audio, reducing the dynamic range, which causes the voice to lose nuances and sound more uniform.
For professionals who use their voice in their work, such as singers, there are specific methods and exercises that help adjust intonation, aiming to bring the voice heard by the audience closer to the one the individual hears themselves.
