Paper List

Bioinformatics

SpikGPT: A High-Accuracy and Interpretable Spiking Attention Framework for Single-Cell Annotation

2025-12-02

This paper addresses the core challenge of robust single-cell annotation across heterogeneous datasets with batch effects and the critical need to ide...
Bioinformatics

Unlocking hidden biomolecular conformational landscapes in diffusion models at inference time

2025-12-02

This paper addresses the core challenge of efficiently and accurately sampling the conformational landscape of biomolecules from diffusion-based struc...
Computational Neuroscience

Personalized optimization of pediatric HD-tDCS for dose consistency and target engagement

2025-12-01

This paper addresses the critical limitation of one-size-fits-all HD-tDCS protocols in pediatric populations by developing a personalized optimization...
Computational Biophysics

Realistic Transition Paths for Large Biomolecular Systems: A Langevin Bridge Approach

2025-12-01

This paper addresses the core challenge of generating physically realistic and computationally efficient transition paths between distinct protein con...
Bioinformatics

Consistent Synthetic Sequences Unlock Structural Diversity in Fully Atomistic De Novo Protein Design

2025-12-01

This paper addresses the core pain point of low sequence-structure alignment in existing synthetic datasets (e.g., AFDB), which severely limits the pe...
Bioinformatics

MoRSAIK: Sequence Motif Reactor Simulation, Analysis and Inference Kit in Python

2025-12-01

This work addresses the computational bottleneck in simulating prebiotic RNA reactor dynamics by developing a Python package that tracks sequence moti...
Bioinformatics

On the Approximation of Phylogenetic Distance Functions by Artificial Neural Networks

2025-12-01

This paper addresses the core challenge of developing computationally efficient and scalable neural network architectures that can learn accurate phyl...
Bioinformatics

EcoCast: A Spatio-Temporal Model for Continual Biodiversity and Climate Risk Forecasting

2025-12-01

This paper addresses the critical bottleneck in conservation: the lack of timely, high-resolution, near-term forecasts of species distribution shifts ...

15 / 18

期刊: ArXiv Preprint

发布日期: 2026-03-14

Computer VisionComputational Neuroscience

Human-like Object Grouping in Self-supervised Vision Transformers

Zuckerman Mind Brain Behavior Institute, Columbia University | Department of Social Science and AI, Hankuk University of Foreign Studies | Nanyang Technological University | University of Hong Kong | Stony Brook University

Hossein Adeli, Seoyoung Ahn, Andrew Luo, Mengmi Zhang, Nikolaus Kriegeskorte, Gregory Zelinsky

30秒速读

IN SHORT: This paper addresses the core challenge of quantifying how well self-supervised vision models capture human-like object grouping in natural scenes, bridging the gap between computational representations and behavioral psychophysics.

核心创新

Methodology Introduces a large-scale behavioral benchmark (1,020 trials) scaling up classical psychophysics to natural images, enabling quantitative comparison between model representations and human object perception.
Methodology Proposes a novel object-centric metric based on ROC analysis of patch-level affinity maps that quantifies object boundary alignment without requiring object-level supervision.
Biology Demonstrates that Gram matrix structure, capturing patch similarity patterns, is a key mechanism driving perceptual alignment between self-supervised models and human vision.

主要结论

Self-supervised Transformer models trained with DINO objectives show strongest alignment with human behavior, with DINOv3 ViT-B achieving 91.9% grouping accuracy and highest noise-normalized Spearman correlation (Fig. 4A).
Object-centric structure in patch representations, quantified by ROC AUC, strongly predicts behavioral alignment across models (correlation shown in Fig. 6B), with DINO-based models consistently outperforming supervised counterparts.
Gram matrix distillation improves supervised models' alignment with human behavior, converging with independent evidence that Gram anchoring enhances DINOv3's feature quality.

研究空白： Previous research has focused on low-level Gestalt cues and simple stimuli, lacking systematic evaluation of how modern vision foundation models align with human object perception in complex natural scenes.

摘要: Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties. However, their alignment with human object perception remains poorly understood. Here, we introduce a behavioral benchmark in which participants make same/different object judgments for dot pairs on naturalistic scenes, scaling up a classical psychophysics paradigm to over 1000 trials. We test a diverse set of vision models using a simple readout from their representations to predict subjects’ reaction times. We observe a steady improvement across model generations, with both architecture and training objective contributing to alignment, and transformer-based models trained with the DINO self-supervised objective showing the strongest performance. To investigate the source of this improvement, we propose a novel metric to quantify the object-centric component of representations by measuring patch similarity within and between objects. Across models, stronger object-centric structure predicts human segmentation behavior more accurately. We further show that matching the Gram matrix of supervised transformer models, capturing similarity structure across image patches, with that of a self-supervised model through distillation improves their alignment with human behavior, converging with the prior finding that Gram anchoring improves DINOv3’s feature quality. Together, these results demonstrate that self-supervised vision models capture object structure in a behaviorally human-like manner, and that Gram matrix structure plays a role in driving perceptual alignment.