Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models".
Jane: Sound symbolism, which describes how people intuitively map speech sounds to perceptual qualities like roundness or sharpness,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, this paper, "Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models," sets out to see if models share our human tendency to map speech sounds onto perceptual qualities like roundness or sharpness. The main idea is that while these models can sometimes match visual shapes, they often struggle with the auditory part of sound symbolism because they miss key acoustic clues.
Jane: Exactly, Tom. The core claim is that the auditory judgments of Speech Language Models diverge from human perception, meaning their internal understanding of how sounds relate to shape isn't quite aligned with how we actually perceive it <ref:2607.10162#pg1>. This matters because it suggests the weakness lies in how speech is represented within these models, not necessarily in their visual capabilities.
Lu: That distinction between auditory and visual representation is what really sparks my imagination; if we can separate those components, we might be able to build more modular systems that handle different sensory inputs more effectively <ref:2607.10162#pg0>.
Meng: I'm thinking about the practical impact on real-world applications; if the auditory cues are missing, how does that translate when we’re building voice assistants or accessibility tools?
Lalam: For me, this points toward a much more nuanced cultural understanding of sound; if an AI can capture these human nuances, it could genuinely improve how we interact with technology on a deeper level <ref:2607.10162#pg0>.
Conclusion: Tom: Looking at the title, "Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models," it really frames the whole discussion around whether AI can genuinely hear sounds the same way we do when we judge shapes based on sound. The authors are Yun-Shao Tsai, Chun-Wei Chen, Chee-En Yu, Yi-Cheng Lin, and Hung-yi Lee from National Taiwan University Taipei.
Jane: What this paper concludes is that Speech Language Models read the visual structure of sound symbolism much like people do when it comes to shape perception. However, their auditory judgments are weaker because they fail to capture the specific acoustic cues that guide human intuition when listening to speech <ref:2607.10162#pg2>.
Lu: The implication here is that building perceptually aligned systems isn't just about having good visual data; it fundamentally requires developing audio representations that accurately mirror the acoustic properties humans rely on for sound symbolism <ref:2607.10162#pg0>. That’s a big conceptual hurdle.
Meng: So, if we take this as a practical directive, it means our focus shifts from just feeding more text to focusing intensely on how we encode the acoustic reality of speech into the model's internal structure <ref:2607.10162#pg1>. I see a clear path for improving voice interaction design if we focus there.
Lalam: I agree with Meng; this suggests that future advances in AI culture will come from making those auditory representations richer, ensuring the AI isn't just mimicking the output but understanding the underlying human acoustic logic <ref:2607.10162#pg0>.
Tom: So, to wrap up, this paper tells us that we can achieve strong visual alignment in SLMs, but we still have a significant gap on the auditory side where they miss those subtle sounds that make sound symbolism work for us. We'll be keeping an eye on how researchers tackle those acoustic cues next.
Graduate Institute of Communication Engineering, National Taiwan University · Graduate Institute of Electrical Engineering, National Taiwan University
eess.AS, cs.CL
Submitted: 2026-07-11
Updated: 2026-10-06
Comments: SLT 2026
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 82/100
The gist: Sound symbolism, which describes how people intuitively map speech sounds to perceptual qualities like roundness or sharpness, is being investigated in Speech Language Models (SLMs) to determine if
Key concepts
- Sound Symbolism
- This refers to the human tendency to intuitively link specific speech sounds to perceptual qualities, such as judging a sound as 'round' or 'sharp.' It explores how our brains naturally connect auditory input with physical shapes and textures without explicit instruction.
- Spectral Tilt
- This is a specific acoustic cue—the way the energy of a sound is distributed across different frequencies. The paper found that humans rely heavily on this feature to judge sound symbolism, whereas current SLMs often fail to utilize it effectively in their auditory judgments.
- Crossmodal Matching
- This involves testing if an AI can correctly link a sound (audio) to a corresponding visual concept (shape). When models use both audio and image inputs, they perform well at matching sounds to shapes, showing that the ability to connect modalities is present in advanced systems.
- Modality Gap
- This describes the difference in performance between how a model processes information from different sensory channels. The study identified a gap where models are strong at visual perception but lack the necessary auditory representation to make accurate sound-based judgments.
Terminology
Summary
Sound symbolism, which describes how people intuitively map speech sounds to perceptual qualities like roundness or sharpness, is being investigated in Speech Language Models (SLMs) to determine if they share this human tendency. The core finding is that while SLMs can sometimes match visual shape perceptions, their auditory judgments often fail to align with human perception because they miss critical acoustic cues, suggesting the weakness lies in how speech is represented rather than vision itself.
The Core Research Questions
The study formulated four research questions to systematically probe this phenomenon:
-
Do SLMs judge speech sounds as round or pointed the way humans do?
-
Are SLMs’ sound-symbolic judgments driven by the same acoustic cues that drive human judgments, such as spectral tilt?
-
Can SLMs match a heard sound to the corresponding shape?
-
If matching fails, can the failure be attributed to the auditory side rather than the visual one?
Experimental Design and Methodology
The researchers conducted four experiments designed to isolate these components of sound symbolism:
-
Experiment 1 (Auditory Channel): This involved a forced-choice task where models chose between a rounded or pointed shape based only on the spoken pseudoword, testing whether they judge sounds as
round or pointed.
-
Experiment 2 (Graded Rating): This experiment used a graded rating scale to ask models to rate pseudowords on both
roundedness
andpointedness,
allowing for the investigation of which acoustic cues drive these judgments. -
Experiment 3 (Crossmodal Matching): This tested whether omni-modal models could match a heard sound to its corresponding shape, which is the closest task to human performance in the bouba/kiki experiment.
-
Experiment 4 (Visual-Only Ablation): This served as a control, removing audio input so models could rate shapes alone, testing whether matching failures are due to auditory limitations or visual perception issues.
Key Findings on Auditory Alignment
The results revealed that SLMs’ auditory judgments diverge from human perception.
Specifically:
'on the auditory side, the models’ soundsymbolic judgments diverge from human perception: the best model reaches barely half of the human ceiling.'
Models failed to exploit acoustic cues that drive human intuition, such as spectral tilt. The analysis showed that for humans, spectral tilt dominates (ρ =.431), with the other cues far behind.
Furthermore, open-weight models fail specifically at the auditory-visual integration that is central to human sound symbolism.
The Role of Visual Perception and Crossmodal Links
In contrast to their auditory weakness, SLMs demonstrated intact shape perception when visual input was present.
'On the visual side, by contrast, omni models (those that accept both audio and image inputs) closely agree with human shape ratings, demonstrating that they can represent the rounded-pointed distinction with high fidelity when the input is visual rather than acoustic.'
Experiment 4 confirmed this: shape perception is intact,
as model correlations against human means ranged from r =.94 to.97 against a human ceiling of.99.
This localization suggests that perceptual alignment depends not on stronger vision but on speech representations that capture the cues humans hear.
Localization of Failure and Representation
The study localized the source of failure through representational similarity analysis (RSA) and logit-lens analysis.
'Perceptual decisions form only in the deepest layers.'
For Qwen3-Omni, perceptual decisions form only in the deepest layers,
with early argmax trajectories showing a diffuse
distribution. The findings indicate a modality gap; while the model lacks the strong ability to separate rounded and pointed in audio, it successfully captures these properties and aligns with the human ratings in the visual domain.
Conclusion on Model Capabilities
The paper concludes that SLMs read the visual structure of sound symbolism much like people do, but their auditory judgments are weak because they miss the acoustic cues that drive human intuitions.
The crossmodal link is not entirely broken; a frontier model succeeded at matching, though this success was plausibly via lexical rather than acoustic routes.
The ultimate requirement for building perceptually aligned systems is developing audio representations that capture the acoustic cues humans rely on.
Index Terms
Sound Symbolism, Speech Language Model, Crossmodal Correspondence, Perceptual Alignment.
The gist:
SLMs’ auditory judgments align poorly with human perception and miss the acoustic cues that drive human intuitions, while their shape perception remains intact in visual inputs.
Improvements for AI systems
Here are specific, actionable improvements for AI systems based on the findings of this paper:
) Improved System Capability: Perceptually Aligned Creative Generation Systems (SLMs)
The primary improvement is shifting Speech Language Models (SLMs) from generating outputs that are merely syntactically correct to generating outputs that align with human perceptual intuitions regarding sound and shape.
-
Acoustic-to-Shape Mapping in Text-to-Image Generation:
-
Improved System Capability: Contextually Perceptual Design Agents
The improved system can perform the following specific functions:
-
Generate images or 3D models based on spoken descriptions that adhere to sound symbolism rules (e.g., generating a sharp, angular robot when prompted with a name like
bomu,
instead of a round one). -
Perform crossmodal matching between auditory input and visual output with high fidelity, moving beyond simple lexical association to genuine acoustic-perceptual correspondence.
-
Develop more robust internal representations that explicitly encode low-level acoustic cues (like spectral tilt) into high-level perceptual judgments, rather than relying solely on superficial language patterns or visual features.
) System Enhancement Strategy: Addressing the Modality Gap and Auditory Weakness
To achieve the above capabilities, focus development on mitigating the identified weaknesses in current SLMs:
-
Representation Learning for Acoustic Cues: Implement novel training objectives or fine-tuning strategies that explicitly force models to learn and utilize acoustic features (e.g., spectral tilt, formant structure) as predictors for perceptual qualities (roundness/sharpness), rather than relying on implicit correlations learned from text/image data alone.
-
Deep Layer Alignment Training: Utilize interpretability tools like the logit lens to identify and target specific layers within the SLM architecture where auditory representations begin to align with visual or human cognitive categories, focusing training efforts there to overcome the observed late divergence in open-weight models.
-
Omni-Modal Integration Refinement: For systems designed for multimodal interaction, prioritize architectures (like those found in Qwen3-Omni) that allow for seamless, high-fidelity integration of audio and visual streams to leverage the model's intact shape perception while simultaneously training the auditory component to match human acoustic intuition.
) Specific Technical Interventions Based on Experiment Findings:
-
For Auditory Alignment (RQ 1 & RQ 2): Train models using contrastive learning objectives that specifically reward representations that correlate with known acoustic cues (e.g., spectral tilt) when mapping sounds to target perceptual categories, rather than just rewarding correct sound classification.
-
For Crossmodal Matching (RQ 3): Develop specialized fine-tuning datasets that explicitly pair pseudowords with their corresponding shapes, forcing the model to learn the direct mapping between the auditory signal and the visual geometry, bypassing potential reliance on intermediate lexical nodes.
-
For Localized Failure Analysis (RQ 4): Implement rigorous ablation studies (as performed in Experiment 4) to quantify exactly where failure occurs—auditory judgment vs. shape perception—to guide targeted architectural fixes for modality-specific weaknesses.
Abstract
Sound symbolism, the human tendency to map speech sounds to perceptual qualities such as roundness or sharpness, arises primarily from the acoustics of speech rather than spelling. Whether Speech Language Models (SLMs) share this tendency remains open, as prior evaluations rely on text or images rather than real speech. We study it using genuine human speech recordings, comparing model judgments against human data across the auditory, crossmodal, and visual components of the effect. We find that SLMs' auditory judgments align poorly with human perception and miss the acoustic cues, such as spectral tilt, that drive human intuitions, and open-weight models cannot reliably link a heard sound to its corresponding shape. With a visual-only control ruling out shape perception, the weakness localizes to how speech is represented, suggesting that perceptual alignment depends not on stronger vision but on speech representations that capture the cues humans hear.
Sources
- Step-Audio 2 Technical Report
- Kimi-Audio Technical Report
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- Qwen3-Omni Technical Report
- Evidence for systematic semantic structure in individual letters
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
- MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions