How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space
Corey D. C. Heath
Abstract
Aesthetics are an important part of the symbolism of artistic works. Although subjective, humans categorize art based on the emotion evoked regardless of modality. What remains under-explored is how AI models form their own aesthetic categorization of human-produced media without explicit labels or cross-modal supervision. We present a self-supervised framework that projects four modalities (text, audio, image and video) into a shared 256-dimensional embedding space and applies iterative clustering to discover aesthetic structure. We discuss the divergence between AI-generated cluster assignments and human affective register labels on a weakly supervised multimodal dataset. This work has applications in understanding how AI structures cross-modal similarity, organizing heterogeneous media collections for Retrieval-Augmented Generation (RAG), and automated data labeling.
Create a lesson
Related papers
Self-Reflective Multi-modal Reasoning for Short-Video Fake News Detection
Pinjie Xu, Yuzhou Yang, Zhikai Tan et al.
Emotion Understanding in Streaming Video with Trajectory-Aware Reliability
Qingsong Wang, Qigong Lei, Zitong Wang et al.
WaveOp-LiteFM: Lightweight Neural-Operator Flow Matching for Satellite-to-Radar Precipitation Retrieval
Chunlei Shi, Yecheng Zhang, Yufeng Zhu et al.
Learning to Prefer Reliably: Error-Augmented Emotion Preference Optimization with Calibrated Fusion
Zilong Huang, Junyi Peng, Junjie Li et al.
EVEREST:Endogenous Vision-Language Reinforcement Reasoning Exploration for Urban Socio-Semantic Segmentation
Qixiu Li, Zhongzhi He, Xiang Zhu et al.
Task-disentangled Low-Rank Adaptation for Versatile Audio-visual Multi-modal Learning Tasks within a Unified Framework
Hanyu Xuan, Mengqi Zhang, Junjun Mao et al.