Human versus Computer Vision
Elena Sirotkina
Abstract
Computer vision saliency models predict where people will look, one map per image, and a billion-dollar predicted-attention industry sells those maps in place of measuring real viewers. I test the leading models from the audience side, against 11.4 million webcam gaze points from 3,023 US adults recruited to national quotas, viewing circulating news photographs. I show that an untrained central marker outperforms every trained network, because the content the networks add on top of the center falls where these audiences never look. What accuracy remains is systematically biased, favoring younger, White, and moderate viewers over older, Black, and ideologically extreme ones. I propose a way forward and build on what a group's own gaze reveals about whether a model can learn that group at all, and I apply it across every demographic axis this sample supports. Ultimately, I show how systems that decide what people see can learn to see everyone, and this study supplies the standard by which such a claim should be judged.
Create a lesson
Related papers
PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image
Sneha Paul, Guile Wu, Bingbing Liu et al.
NormLift: From Lifted Features To Semantic Reliability In 3D Gaussian Splatting
Yihan Zang, Da Li, Dominik Engel et al.
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
Girish A. Koushik, Diptesh Kanojia, Helen Treharne
Copy What Is Seen, Generate What Is Not: Training-Free Anomaly-Aware Video Restoration
Zhida Qu, Shengchao Chen
Using OCR Heads to Verbalize Image Semantics
Sheridan Feucht, Benno Krojer, Sarah Wang et al.
DISTA-Net++: Rethinking Infrared Small Target Unmixing Beyond Sub-Pixel Separation
Mengze Xu, Zhu Liu, Weidong Sheng et al.