Explicitly Modeling Subcortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness

Lucas Piper; Arlindo L. Oliveira; Tiago Marques

Explicitly Modeling Subcortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness

Lucas Piper, Arlindo L. Oliveira, Tiago Marques

TL;DR

CNN robustness gaps relative to biological vision are addressed by EVNets, which explicitly model subcortical vision through a fixed SubcorticalBlock paired with a VOneBlock front end before a CNN backend. The front-end cascade implements a center–surround DoG with light adaptation and contrast normalization, plus a noise generator, forming a fixed, neuro-inspired early-vision pipeline with a $7^ ^ ightarrow$FoV and extended spatial-frequency coverage via the Gabor filter bank. EVNets achieve stronger V1 alignment (including extra-classical receptive-field properties) and a substantial 9.3% gain in a robustness score over the baseline CNN, with additional additive gains when combined with PRIME data augmentation. The improvements generalize across back-ends and arise from complementary architectural priors and training-based strategies, suggesting a practical path to more robust, brain-aligned vision systems.

Abstract

Convolutional neural networks (CNNs) trained on object recognition achieve high task performance but continue to exhibit vulnerability under a range of visual perturbations and out-of-domain images, when compared with biological vision. Prior work has demonstrated that coupling a standard CNN with a front-end (VOneBlock) that mimics the primate primary visual cortex (V1) can improve overall model robustness. Expanding on this, we introduce Early Vision Networks (EVNets), a new class of hybrid CNNs that combine the VOneBlock with a novel SubcorticalBlock, whose architecture draws from computational models in neuroscience and is parameterized to maximize alignment with subcortical responses reported across multiple experimental studies. Without being optimized to do so, the assembly of the SubcorticalBlock with the VOneBlock improved V1 alignment across most standard V1 benchmarks, and better modeled extra-classical receptive field phenomena. In addition, EVNets exhibit stronger emergent shape bias and outperform the base CNN architecture by 9.3% on an aggregate benchmark of robustness evaluations, including adversarial perturbations, common corruptions, and domain shifts. Finally, we show that EVNets can be further improved when paired with a state-of-the-art data augmentation technique, surpassing the performance of the isolated data augmentation approach by 6.2% on our robustness benchmark. This result reveals complementary benefits between changes in architecture to better mimic biology and training-based machine learning approaches.

Explicitly Modeling Subcortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness

TL;DR

Abstract

Explicitly Modeling Subcortical Vision with a Neuro-Inspired Front-End Improves CNN Robustness

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (8)