Sep 22 – 24, 2026
SISSA
Europe/Rome timezone

Texture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human Perception

Sep 23, 2026, 3:45 PM
2h 15m
Aula Magna “Paolo Budinich” (SISSA)

Aula Magna “Paolo Budinich”

SISSA

Via Bonomea 265, Trieste

Speaker

Ludovica de Paolis (SISSA)

Description

In computational vision science, Convolutional Neural Networks (CNNs) have
emerged as a popular model of biological vision because of the alignment they
can exhibit with neural and behavioral data in humans and animals. However, it
is unclear to what extent this alignment persists for visual tasks that stray from
the canonical object recognition with well-defined semantic content. In this study,
we diverge from the common object-centric view by focusing on another aspect of
vision: texture perception. We consider textures of different complexity generated
with three different algorithms from the same source images. Using a rank-based
statistic, we quantify the information encoded in the internal representations of
a CNN and three Vision Transformers (ViTs), and we compare the similarity of
these representations to those inferred from human psychophysics data. We find
that the representation of textures is aligned in different ViTs, but not between
the ViTs and the CNN; that ViTs form similar representations for textures of dif-
ferent complexity; that human performance in recognizing textures can be better
predicted from ViTs representations rather than CNN representations. Taken to-
gether, these results suggest that ViTs may capture more faithfully than CNNs
how texture patterns are visually processed by humans, and that the representa-
tional geometry of texture stimuli in computational models may be driven by the
network architecture.

Preferred Presentation Oral Presentation

Author

Ludovica de Paolis (SISSA)

Co-authors

Prof. Alessandro Laio (SISSA) Prof. Marco Baroni (Universitat Pompeu Fabra, ICREA) Prof. Eugenio Piasini (SISSA)

Presentation materials

There are no materials yet.