Uncovering Semantic Selectivity of Latent Groups in Higher Visual Cortex with Mutual Information-Guided Diffusion
Yule Wang, Joseph Yu, Chengrui Li, Weihan Li, Anqi Wu
TL;DR
The paper tackles how object-centered visual information is represented in higher visual cortex by introducing MIG-Vis, which combines a group-wise disentangled VAE to extract neural latent groups with a mutual-information-guided diffusion framework to visualize their semantic attributes. It demonstrates that latent groups in IT cortex encode distinct semantic factors, such as intra-category pose, inter-category semantics, and intra-category content details, and validates these findings through deterministic DDIM-based editing guided by MI. The approach yields high neural reconstruction quality and significantly higher latent disentanglement than standard VAEs, with robust quantitative and qualitative evidence from macaque IT data. This work provides direct interpretable insight into the structured, multi-dimensional nature of visual coding in primate cortex and offers a reproducible tool for exploring neural subspace geometry.
Abstract
Understanding how neural populations in higher visual areas encode object-centered visual information remains a central challenge in computational neuroscience. Prior works have investigated representational alignment between artificial neural networks and the visual cortex. Nevertheless, these findings are indirect and offer limited insights to the structure of neural populations themselves. Similarly, decoding-based methods have quantified semantic features from neural populations but have not uncovered their underlying organizations. This leaves open a scientific question: "how feature-specific visual information is distributed across neural populations in higher visual areas, and whether it is organized into structured, semantically meaningful subspaces." To tackle this problem, we present MIG-Vis, a method that leverages the generative power of diffusion models to visualize and validate the visual-semantic attributes encoded in neural latent subspaces. Our method first uses a variational autoencoder to infer a group-wise disentangled neural latent subspace from neural populations. Subsequently, we propose a mutual information (MI)-guided diffusion synthesis procedure to visualize the specific visual-semantic features encoded by each latent group. We validate MIG-Vis on multi-session neural spiking datasets from the inferior temporal (IT) cortex of two macaques. The synthesized results demonstrate that our method identifies neural latent groups with clear semantic selectivity to diverse visual features, including object pose, inter-category transformations, and intra-class content. These findings provide direct, interpretable evidence of structured semantic representation in the higher visual cortex and advance our understanding of its encoding principles.
