Pseudo-Likelihood Ratio Screening based on Network Data with Applications

Abstract

Social network platforms today generate vast amounts of data, including network structures and a large number of user-defined tags, which reflect users' interests. The dimensionality of these personalized tags can be ultra-high, posing challenges for model analysis in targeted preference analysis. Traditional categorical feature screening methods overlook the network structure, which can lead to incorrect feature set and suboptimal prediction accuracy. This study focuses on feature screening for network-involved preference analysis based on ultra-high-dimensional categorical tags. We introduce the concepts of self-related features and network-related features, defined as those directly related to the response and those related to the network structure, respectively. We then propose a pseudo-likelihood ratio feature screening procedure that identifies both types of features. Theoretical properties of this procedure under different scenarios are thoroughly investigated. Extensive simulations and real data analysis on Sina Weibo validate our findings.

0

Turn this paper into a full lesson

ArcXiv compiles a staged curriculum from this paper: 8-12 lessons across beginner → advanced, synthesised section guides, visuals, flashcards, a quiz, exercises, and on-demand deep dives per section. Grounded in the abstract, never invented.

Discussion (0)

Sign in to join the discussion.

Loading comments…