Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?
Michael Aerni, Joshua Swanson, Kristina Nikolić, Florian Tramèr
TL;DR
Modal aphasia reveals a robust cross-modal dissociation in unified multimodal models: they can visually reproduce learned concepts with high fidelity but struggle to articulate the same concepts in text. Through frontier-model demonstrations and open-weight controlled experiments, the work shows a consistent image-text gap with text descriptions exhibiting substantially more hallucinations and lower accuracy. The findings persist across synthetic-face and abstract-concept datasets, suggesting a fundamental limitation in how current architectures store and transfer cross-modal knowledge. A safety case study further indicates that text-only safeguards may fail to block unsafe concepts that remain accessible via other modalities, underscoring the need for models that integrate visualization into reasoning or adopt cross-modal alignment strategies. Overall, the work highlights a core challenge in achieving truly unified multimodal understanding and provides reproducible experimental resources to probe cross-modal knowledge transfer.
Abstract
We present modal aphasia, a systematic dissociation in which current unified multimodal models accurately memorize concepts visually but fail to articulate them in writing, despite being trained on images and text simultaneously. For one, we show that leading frontier models can generate near-perfect reproductions of iconic movie artwork, but confuse crucial details when asked for textual descriptions. We corroborate those findings through controlled experiments on synthetic datasets in multiple architectures. Our experiments confirm that modal aphasia reliably emerges as a fundamental property of current unified multimodal models, not just as a training artifact. In practice, modal aphasia can introduce vulnerabilities in AI safety frameworks, as safeguards applied to one modality may leave harmful concepts accessible in other modalities. We demonstrate this risk by showing how a model aligned solely on text remains capable of generating unsafe images.
