May 2025 arXiv papers — page 125
Showing 12,401–12,500 of 24,552 papers
Yinzhe Wang, Yiwen Xiao, Hu Wang, Yiping Xu
Multi-view stereo (MVS) models based on progressive depth hypothesis narrowing have made remarkable advancements. However, existing methods haven't fully utilized the potential that the depth coverage of individual instances is smaller than that of the entire scene, which restricts further improvements in depth estimation precision. Moreover, inevitable devi
Subhayan Saha, Giovanni Barbarino, Nicolas Gillis
Tensor decompositions have become a central tool in data science, with applications in areas such as data analysis, signal processing, and machine learning. A key property of many tensor decompositions, such as the canonical polyadic decomposition, is identifiability: the factors are unique, up to trivial scaling and permutation ambiguities. This allows one
Shogo Kusano, Masayuki Uchida
We study structural equation modeling (SEM) for diffusion processes with jumps. Based on high-frequency data, we consider the parameter estimation and the goodness-of-fit test in the SEM. Using a threshold method, we propose the quasi-likelihood of the SEM and prove that the quasi-maximum likelihood estimator has consistency and asymptotic normality. To exam
Qichen Sun, Zhengrui Guo, Rui Peng, Hao Chen
Recent advances in computational pathology and artificial intelligence have significantly enhanced the utilization of gigapixel whole-slide images and and additional modalities (e.g., genomics) for pathological diagnosis. Although deep learning has demonstrated strong potential in pathology, several key challenges persist: (1) fusing heterogeneous data types
Confidence-Regulated Generative Diffusion Models for Reliable AI Agent Migration in Vehicular Metaverses
cs.LGYingkai Kang, Jiawen Kang, Jinbo Wen, Tao Zhang
Vehicular metaverses are an emerging paradigm that merges intelligent transportation systems with virtual spaces, leveraging advanced digital twin and Artificial Intelligence (AI) technologies to seamlessly integrate vehicles, users, and digital environments. In this paradigm, vehicular AI agents are endowed with environment perception, decision-making, and
Serge Dolgikh
This study explores the emergence of counter-inferential behavior in natural and artificial cognitive systems, that is, patterns in which agents misattribute empirical success or suppress adaptation, leading to epistemic rigidity or maladaptive stability. We analyze archetypal scenarios in which such behavior arises: reinforcement of stability through reward
Zhichen Zeng, Ruizhong Qiu, Wenxuan Bao, Tianxin Wei
Graph neural networks, despite their impressive performance, are highly vulnerable to distribution shifts on graphs. Existing graph domain adaptation (graph DA) methods often implicitly assume a mild shift between source and target graphs, limiting their applicability to real-world scenarios with large shifts. Gradual domain adaptation (GDA) has emerged as a
Yingchen He, Christian D. Weilbach, Martyna E. Wojciechowska, Yuxuan Zhang
Advances in deep generative modeling have made it increasingly plausible to train human-level embodied agents. Yet progress has been limited by the absence of large-scale, real-time, multi-modal, and socially interactive datasets that reflect the sensory-motor complexity of natural environments. To address this, we present PLAICraft, a novel data collection
Efficient computation of complementary set partitions, with applications to an extension and estimation of generalized cumulants
math.STElvira Di Nardo, Giuseppe Guarino
This paper develops new combinatorial approaches to analyze and compute special set partitions, called complementary set partitions, which are fundamental in the study of generalized cumulants. Moving away from traditional graph-based and algebraic methods, a simple and fast algorithm is proposed to list complementary set partitions based on two-block partit
Joel Jang, Seonghyeon Ye, Zongyu Lin, Jiannan Xiang
We introduce DreamGen, a simple yet highly effective 4-stage pipeline for training robot policies that generalize across behaviors and environments through neural trajectories - synthetic robot data generated from video world models. DreamGen leverages state-of-the-art image-to-video generative models, adapting them to the target robot embodiment to produce
Itinerant ferromagnetism in an SU(3) Fermi-Hubbard model at finite temperatures: A dynamical mean-field theory study
cond-mat.str-elJuntaro Fujii, Kazuki Yamamoto, Akihisa Koga
We investigate an SU(3) Fermi-Hubbard model on a hypercubic lattice at finite temperatures, combining dynamical mean-field theory with continuous-time quantum Monte Carlo simulations. Taking strong correlations into account carefully, we find a ferromagnetically ordered state, in which one of the three components becomes dominant, when holes are doped away f
Jiabin Chen, Haiping Wang, Jinpeng Li, Yuan Liu
We propose SpatialLLM, a novel approach advancing spatial intelligence tasks in complex urban scenes. Unlike previous methods requiring geographic analysis tools or domain expertise, SpatialLLM is a unified language model directly addressing various spatial intelligence tasks without any training, fine-tuning, or expert intervention. The core of SpatialLLM l
Tianming Liang, Haichao Jiang, Yuting Yang, Chaolei Tan
Referring video object segmentation (RVOS) aims to identify, track and segment the objects in a video based on language descriptions, which has received great attention in recent years. However, existing datasets remain focus on short video clips within several seconds, with salient objects visible in most frames. To advance the task towards more practical s
Ke Yang, Kevin Ros, Shankar Kumar Senthil Kumar, ChengXiang Zhai
Just-in-time Information Recommendation (JIR) is a service designed to deliver the most relevant information precisely when users need it, , addressing their knowledge gaps with minimal effort and boosting decision-making and efficiency in daily life. Advances in device-efficient deployment of foundation models and the growing use of intelligent wearable dev
Re-experiment Smart: a Novel Method to Enhance Data-driven Prediction of Mechanical Properties of Epoxy Polymers
cond-mat.softWanshan Cui, Yejin Jeong, Inwook Song, Gyuri Kim
Accurate prediction of polymer material properties through data-driven approaches greatly accelerates novel material development by reducing redundant experiments and trial-and-error processes. However, inevitable outliers in empirical measurements can severely skew machine learning results, leading to erroneous prediction models and suboptimal material desi
Shuyang Dong, Shangtong Zhang, Lu Feng
Reinforcement Learning (RL) has shown great promise in domains like healthcare and robotics but often struggles with adoption due to its lack of interpretability. Counterfactual explanations, which address "what if" scenarios, provide a promising avenue for understanding RL decisions but remain underexplored for continuous action spaces. We propose a novel a
Utsav Banerjee, Chiraag Juvekar, Yong Ki Lee, Leibo Liu
Security is increasingly more important in designing chips and systems based on them, and the International Solid-State Circuits Conference (ISSCC), the leading conference for presenting advances in solid-state circuits and semiconductor technology, is committed to hardware security by establishing the security subcommittee since 2024. In the past two years,
Sushmita Gupta, Pallavi Jain, Souvik Saha, Saket Saurabh
Multiwinner Elections have emerged as a prominent area of research with numerous practical applications. We contribute to this area by designing parameterized approximation algorithms and also resolving an open question by Yang and Wang [AAMAS'18]. More formally, given a set of candidates, \mathcal{C}, a set of voters,\mathcal{V}, approving a subset of candi
Yixin Chen, Xiaoyang Wang, Wanghui Li, Mohan Chen
Tin (Sn) plays a crucial role in studying the dynamic mechanical responses of ductile metals under shock loading. Atomistic simulations serves to unveil the nano-scale mechanisms for critical behaviors of dynamic responses. However, existing empirical potentials for Sn often lack sufficient accuracy when applied in such simulation. Particularly, the solid-so
Chaofan Li, Jianlyu Chen, Yingxia Shao, Defu Lian
Code embedding models attract increasing attention due to the widespread popularity of retrieval-augmented generation (RAG) in software development. These models are expected to capture the rich semantic relationships inherent to code, which differ significantly from those found in text. However, existing models remain severely limited due to the scarcity of
Wenqi Tong, H. Alaeian, F. Robicheaux
We study the steady-state behavior of the open Dicke model, which describes the collective interaction of $N$ spin-$1/2$ particles with a lossy, quantized cavity mode and exhibits a superradiant phase transition above a critical light-matter coupling. While the standard model conserves total spin, Kirton and Keeling \cite{PhysRevLett.118.123602} demonstrated
Wei Hu, Danyang Huang, Bo Zhang
Social network platforms today generate vast amounts of data, including network structures and a large number of user-defined tags, which reflect users' interests. The dimensionality of these personalized tags can be ultra-high, posing challenges for model analysis in targeted preference analysis. Traditional categorical feature screening methods overlook th
Kenya Abe, Kunihiro Takeoka, Makoto P. Kato, Masafumi Oyamada
Query expansion (QE) enhances retrieval by incorporating relevant terms, with large language models (LLMs) offering an effective alternative to traditional rule-based and statistical methods. However, LLM-based QE suffers from a fundamental limitation: it often fails to generate relevant knowledge, degrading search performance. Prior studies have focused on
Luyao Lei, Shuo Xu, Yifan Bai, Xing Wei
The performance of multi-modal 3D occupancy prediction is limited by ineffective fusion, mainly due to geometry-semantics mismatch from fixed fusion strategies and surface detail loss caused by sparse, noisy annotations. The mismatch stems from the heterogeneous scale and distribution of point cloud and image features, leading to biased matching under fixed
TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion
cs.ROKhang Nguyen, Khai Nguyen, An T. Le, Jan Peters
Robot learning in high-dimensional control settings, such as humanoid locomotion, presents persistent challenges for reinforcement learning (RL) algorithms due to unstable dynamics, complex contact interactions, and sensitivity to distributional shifts during training. Model-based methods, \textit{e.g.}, Temporal-Difference Model Predictive Control (TD-MPC),
Ziwei Xu, Udit Sanghi, Mohan Kankanhalli
Large Language Models (LLMs) are increasingly deployed in interactions where they are prompted to adopt personas. This paper investigates whether such persona conditioning affects model safety under bullying, an adversarial manipulation that applies psychological pressures in order to force the victim to comply to the attacker. We introduce a simulation fram
Haojie Hou, Yan-Xia Ren, Renming Song
Let $\{(X_t)_{t\geq 0}, \mathbb{P}_{\delta_x}, x\in E\}$ be a supercritical branching Markov process (which is not necessary symmetric) on a locally compact metric measure space $(E,\mu)$ with spatially dependent local branching mechanism. Under some assumptions on the semigroup of the spatial motion, we first prove law of iterated logarithm type results for
Kian Kai Ang, Guy Farrelly, Cheryl Pope, Damith C. Ranasinghe
We develop QUICtester, an automated approach for uncovering non-compliant behaviors in the ratified QUIC protocol implementations (RFC 9000/9001). QUICtester leverages active automata learning to abstract the behavior of a QUIC implementation into a finite state machine (FSM) representation. Unlike prior noncompliance checking methods, to help uncover state
Formation of a Possible Black-hole Ultracompact X-ray Binary with the Shortest Orbital Period
astro-ph.HEXing-Peng Yang, Kun Xu, Zhi-Fu Gao, Long Jiang
In the bulge of M31, the Chandra observations discovered a possible black hole (BH) ultracompact X-ray binary (UCXB) Seq.1 with an orbital period of 7.7 minutes and a maximum X-ray luminosity $L_{\rm X}=1.09^{+0.02}_{-0.01}\times10^{38}~ \rm erg\,s^{-1}$ in the $0.5-8$ keV band. The minimum orbital period of the BH UCXBs predicted by the standard magnetic br
Arjun Ramesh Kaushik, Bharat Chandra Yalavarthi, Arun Ross, Vishnu Boddeti
In today's data-driven analytics landscape, deep learning has become a powerful tool, with latent representations, known as embeddings, playing a central role in several applications. In the face analytics domain, such embeddings are commonly used for biometric recognition (e.g., face identification). However, these embeddings, or templates, can inadvertentl
Li Lai, Jia Li
The Chowla--Milnor conjecture predicts the linear independence of certain Hurwitz zeta values. In this paper, we prove that for any fixed integer $k \geqslant 2$, the dimension of the $\mathbb{Q}$-linear span of $\zeta(k,a/q)-(-1)^{k}\zeta(k,1-a/q)$ ($1 \leqslant a < q/2$, $\gcd(a,q)=1$) is at least $(c -o(1)) \cdot \log q$ as the positive integer $q \to +\i
RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations
cs.LGSeungmin Kim, Sohee Park, Donghyun Kim, Jisu Lee
With the advancement of AI-based speech synthesis technologies such as Deep Voice, there is an increasing risk of voice spoofing attacks, including voice phishing and fake news, through unauthorized use of others' voices. Existing defenses that inject adversarial perturbations directly into audio signals have limited effectiveness, as these perturbations can
Fei Xie, Jiahao Nie, Yujin Tang, Wenkang Zhang
Recent State Space Models (SSM), especially Mamba, have demonstrated impressive performance in visual modeling and possess superior model efficiency. However, the application of Mamba to visual tasks suffers inferior performance due to three main constraints existing in the sequential model: 1) Casual computing is incapable of accessing global context; 2) Lo
Dmitry Nesterov
We present an alternative formulation of generalized unimodular gravity (GUMG), a class of modifications to general relativity characterized by a special partial breaking of general coordinate covariance. The action for this formulation is derived constructively through a sequence of equivalent representations, starting from the original GUMG setup and exten
Yinlin Zhu, Xunkai Li, Jishuo Jia, Miao Hu
Recent advances in graph machine learning have shifted to data-centric paradigms, driven by two emerging fields: (1) Federated graph learning (FGL) enables multi-client collaboration but faces challenges from data and task heterogeneity, limiting its practicality; (2) Graph foundation models (GFM) offer strong domain generalization but are usually trained on
Yihong Huang, Chen Chu
Key feature fields need bigger embedding dimensionality, others need smaller. This demands automated dimension allocation. Existing approaches, such as pruning or Neural Architecture Search (NAS), require training a memory-intensive SuperNet that enumerates all possible dimension combinations, which is infeasible for large feature spaces. We propose DimGrow,
Hana Satou, Alan Mitkiy
Transfer learning across domains with distribution shift remains a fundamental challenge in building robust and adaptable machine learning systems. While adversarial perturbations are traditionally viewed as threats that expose model vulnerabilities, recent studies suggest that they can also serve as constructive tools for data augmentation. In this work, we
Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on Inequalities
cs.AIHaoyu Zhao, Yihan Geng, Shange Tang, Yong Lin
LLM-based formal proof assistants (e.g., in Lean) hold great promise for automating mathematical discovery. But beyond syntactic correctness, do these systems truly understand mathematical structure as humans do? We investigate this question in context of mathematical inequalities -- specifically the prover's ability to recognize that the given problem simpl
Zhuoheng Wang, Jinyin Zhou, Qi Wu
Humanoid soccer dribbling is a highly challenging task that demands dexterous ball manipulation while maintaining dynamic balance. Traditional rule-based methods often struggle to achieve accurate ball control due to their reliance on fixed walking patterns and limited adaptability to real-time ball dynamics. To address these challenges, we propose a two-sta
Measuring the cosmic dipole with golden dark sirens in the era of next-generation ground-based gravitational wave detectors
gr-qcAnson Chen
The tensions between cosmological parameter measurements from the early-universe and the late-universe datasets offer an exciting opportunity to explore new physics, if not accounted for unknown systematics. Apart from the well-known Hubble tension, a tension up to $4.9 \sigma$ in the cosmic dipole has also been reported. While the cosmic dipole is mainly in
Shristi Das Biswas, Arani Roy, Kaushik Roy
As Text-to-Image models continue to evolve, so does the risk of generating unsafe, copyrighted, or privacy-violating content. Existing safety interventions - ranging from training data curation and model fine-tuning to inference-time filtering and guidance - often suffer from incomplete concept removal, susceptibility to jail-breaking, computational ineffici
Xingyu Wang, Mingsen Wang, Wenbo Shen, Rui Chang
As the default package manager for Node.js, npm has become one of the largest package management systems in the world. To facilitate dependency management for developers, npm supports a special type of dependency, Peer Dependency, whose installation and usage differ from regular dependencies. However, conflicts between peer dependencies can trap the npm clie
Won-Young Hwang, Kicheon Kang
We propose an experimental scheme to probe the quantum statistics of two identical particles. The transition between the quantum and classical statistics of two identical particles is described by the particles having identical multiple internal energy levels. We show that effective distinguishability emerges as the thermal energy increases with respect to t
Mingyuan Zhou, Yi Gu, Zhendong Wang
Diffusion distillation has emerged as a promising strategy for accelerating text-to-image (T2I) diffusion models by distilling a pretrained score network into a one- or few-step generator. While existing methods have made notable progress, they often rely on real or teacher-synthesized images to perform well when distilling high-resolution T2I diffusion mode
Zhi-Ming Yang, Huan Li
Nonsymmorphic symmetries can give rise to Dirac semimetal (DSM) states. However, few studies have been conducted on DSMs in interacting systems. Here, we induce interacting DSM states in nonsymmorphic iridium oxides SrIrO$_3$, BaIrO$_3$ and CaIrO$_3$, and contend that the interaction of electron-electron correlations, strong spin-orbital coupling, and symmet
Pengxin Guo, Yinong Wang, Wei Li, Mengting Liu
LLM pruning has emerged as a promising technology for compressing LLMs, enabling their deployment on resource-limited devices. However, current methodologies typically require access to public calibration samples, which can be challenging to obtain in privacy-sensitive domains. To address this issue, we introduce FedPrLLM, a comprehensive federated pruning f
Tonglong Wei, Yan Lin, Zeyu Zhou, Haomin Wen
Vehicle GPS trajectories provide valuable movement information that supports various downstream tasks and applications. A desirable trajectory learning model should be able to transfer across regions and tasks without retraining, avoiding the need to maintain multiple specialized models and subpar performance with limited training data. However, each region
TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning
cs.CVLihong Chen, Hossein Hassani, Soodeh Nikan
Vision-Language Models (VLMs) have shown remarkable potential in advancing autonomous driving by leveraging multi-modal fusion in order to enhance scene perception, reasoning, and decision-making. Despite their potential, existing models suffer from computational overhead and inefficient integration of multi-view sensor data that make them impractical for re
Abhinaba Roy, Geeta Puri, Dorien Herremans
We present Text2midi-InferAlign, a novel technique for improving symbolic music generation at inference time. Our method leverages text-to-audio alignment and music structural alignment rewards during inference to encourage the generated music to be consistent with the input caption. Specifically, we introduce two objectives scores: a text-audio consistency
Decompile-Bench: Million-Scale Binary-Source Function Pairs for Real-World Binary Decompilation
cs.SEHanzhuo Tan, Xiaolong Tian, Hanrui Qi, Jiaming Liu
Recent advances in LLM-based decompilers have been shown effective to convert low-level binaries into human-readable source code. However, there still lacks a comprehensive benchmark that provides large-scale binary-source function pairs, which is critical for advancing the LLM decompilation technology. Creating accurate binary-source mappings incurs severe
Zihan Su, Xuerui Qiu, Hongbin Xu, Tangyu Jiang
The explosive growth of generative video models has amplified the demand for reliable copyright preservation of AI-generated content. Despite its popularity in image synthesis, invisible generative watermarking remains largely underexplored in video generation. To address this gap, we propose Safe-Sora, the first framework to embed graphical watermarks direc
Huimin Xu, Houjiang Liu, Yan Leng, Ying Ding
CSCW has long examined how emerging technologies reshape the ways researchers collaborate and produce knowledge, with scientific knowledge production as a central area of focus. As AI becomes increasingly integrated into scientific research, understanding how researchers adapt to it reveals timely opportunities for CSCW research -- particularly in supporting
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
cs.AIKe Chen, Yufei Zhou, Xitong Zhang, Haohan Wang
Automatic prompt generation plays a crucial role in enabling general-purpose multi-agent systems to perform diverse tasks autonomously. Existing methods typically evaluate prompts based on their immediate task performance, overlooking the intrinsic qualities that determine their reliability. This outcome-centric view not only limits interpretability but also
Ryan Spears, Moonyoung Lee, George Kantor, Oliver Kroemer
Contact-rich manipulation tasks in agriculture, such as pruning and harvesting, require robots to physically interact with tree structures to maneuver through cluttered foliage. Identifying whether the robot is contacting rigid or soft materials is critical for the downstream manipulation policy to be safe, yet vision alone is often insufficient due to occlu
Ziqing Xing, Zhaoyang Zhang, Zirui Chen, Hongning Ruan
In this paper, we incorporate physical knowledge into learning-based high-precision target sensing using the multi-view channel state information (CSI) between multiple base stations (BSs) and user equipment (UEs). Such kind of multi-view sensing problem can be naturally cast into a conditional generation framework. To this end, we design a bipartite neural
Xukai Liu, Ye Liu, Shiwen Wu, Yanghai Zhang
Recent advances in large language models (LLMs) have led to impressive progress in natural language generation, yet their tendency to produce hallucinated or unsubstantiated content remains a critical concern. To improve factual reliability, Retrieval-Augmented Generation (RAG) integrates external knowledge during inference. However, existing RAG systems fac
Digital Twins in the Cloud: A Modular, Scalable and Interoperable Framework for Accelerating Verification and Validation of Autonomous Driving Solutions
cs.ROTanmay Vilas Samak, Chinmay Vilas Samak, Giovanni Martino, Pranav Nair
Verification and validation (V&V) of autonomous vehicles (AVs) typically requires exhaustive testing across a variety of operating environments and driving scenarios including rare, extreme, or hazardous situations that might be difficult or impossible to capture in reality. Additionally, physical V&V methods such as track-based evaluations or public-road te
Ziqi Wen, Jonathan Skaza, Shravan Murlidaran, William Y. Wang
Although models exist that predict human response times (RTs) in tasks such as target search and visual discrimination, the development of image-computable predictors for scene understanding time remains an open challenge. Recent advances in vision-language models (VLMs), which can generate scene descriptions for arbitrary images, combined with the availabil
Jessica Foo, Pradyumna Shyama Prasad, Shaun Khoo
While the capabilities of large language models (LLMs) have progressed significantly, their use in high-stakes applications have been limited due to risks of hallucination. One key approach in reducing hallucination is retrieval-augmented generation (RAG), but even in such setups, LLMs may still hallucinate when presented with questions outside of the knowle
Xianzhe Dong, Tongxuan Liu, Yuting Zeng, Liangyu Liu
Multimodal Large Language Models (MLLMs) have been rapidly advancing, enabling cross-modal understanding and generation, and propelling artificial intelligence towards artificial general intelligence. However, existing MLLM inference systems are typically designed based on the architecture of language models, integrating image processing and language process
Shuang Gao, Peter E. Caines
Transmission Neural Networks (TransNNs) introduced by Gao and Caines (2022) connect virus spread models over networks and neural networks with tuneable activation functions. This paper presents the approximation technique and the underlying assumptions employed by TransNNs in relation to the corresponding Markovian Susceptible-Infected-Susceptible (SIS) mode
Yongchang Gao, Meiling Jin, Zhaofei Yu, Tiejun Huang
Spike cameras offer unique sensing capabilities but their sparse, asynchronous output challenges semantic understanding, especially for Spike Video-Language Alignment (Spike-VLA) where models like CLIP underperform due to modality mismatch. We introduce SPKLIP, the first architecture specifically for Spike-VLA. SPKLIP employs a hierarchical spike feature ext
Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language Models
cs.CRYisheng Zhong, Yizhu Wen, Junfeng Guo, Mehran Kafai
The protection of cyber Intellectual Property (IP) such as web content is an increasingly critical concern. The rise of large language models (LLMs) with online retrieval capabilities enables convenient access to information but often undermines the rights of original content creators. As users increasingly rely on LLM-generated responses, they gradually dim
Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals
cs.CLYuxin Lin, Yinglin Zheng, Ming Zeng, Wangzheng Shi
This paper addresses the gap in predicting turn-taking and backchannel actions in human-machine conversations using multi-modal signals (linguistic, acoustic, and visual). To overcome the limitation of existing datasets, we propose an automatic data collection pipeline that allows us to collect and annotate over 210 hours of human conversation videos. From t
High-dimensional structure underlying individual differences in naturalistic visual experience
q-bio.NCChihye Han, Michael F. Bonner
How do different brains create unique visual experiences from identical sensory input? While neural representations vary across individuals, the fundamental architecture underlying these differences remains poorly understood. Here, we reveal that individual visual experience emerges from a high-dimensional neural geometry across the visual cortical hierarchy
Konstantin Batygin, Fred C. Adams
The formation and early evolution of Jupiter played a pivotal role in sculpting the large-scale architecture of the solar system, intertwining the narrative of Jovian early years with the broader story of the solar system's origins. The details and chronology of Jupiter's formation, however, remain elusive, primarily due to the inherent uncertainties of accr
$\texttt{DIAMONDs}$: A Dataset for $\mathbb{D}$ynamic $\mathbb{I}$nformation $\mathbb{A}$nd $\mathbb{M}$ental modeling $\mathbb{O}$f $\mathbb{N}$umeric $\mathbb{D}$iscussions
cs.AISayontan Ghosh, Mahnaz Koupaee, Yash Kumar Lal, Pegah Alipoormolabashi
Understanding multiparty conversations demands robust Theory of Mind (ToM) capabilities, including the ability to track dynamic information, manage knowledge asymmetries, and distinguish relevant information across extended exchanges. To advance ToM evaluation in such settings, we present a carefully designed scalable methodology for generating high-quality
Unveil Multi-Picture Descriptions for Multilingual Mild Cognitive Impairment Detection via Contrastive Learning
cs.CLKristin Qi, Jiali Cheng, Youxiang Zhu, Hadi Amiri
Detecting Mild Cognitive Impairment from picture descriptions is critical yet challenging, especially in multilingual and multiple picture settings. Prior work has primarily focused on English speakers describing a single picture (e.g., the 'Cookie Theft'). The TAUKDIAL-2024 challenge expands this scope by introducing multilingual speakers and multiple pictu
Karthik Urs, Jessica Carlson, Aditya Srinivas Manohar, Michael Rakowiecki
Robotic models are useful for independently varying specific features, but most quadrupedal robots differ so greatly from animal morphologies that they have minimal biomechanical relevance. Commercially available quadrupedal robots are also prohibitively expensive for biological research programs and difficult to customize. Here, we present a low-cost quadru
SafeMove-RL: A Certifiable Reinforcement Learning Framework for Dynamic Motion Constraints in Trajectory Planning
cs.ROTengfei Liu, Haoyang Zhong, Jiazheng Hu, Tan Zhang
This study presents a dynamic safety margin-based reinforcement learning framework for local motion planning in dynamic and uncertain environments. The proposed planner integrates real-time trajectory optimization with adaptive gap analysis, enabling effective feasibility assessment under partial observability constraints. To address safety-critical computat
Jung Hoon Lee, Sujith Vijayan
Deep learning (DL) is a powerful tool that can solve complex problems, and thus, it seems natural to assume that DL can be used to enhance the security of wireless communication. However, deploying DL models to edge devices in wireless networks is challenging, as they require significant amounts of computing and power resources. Notably, Spiking Neural Netwo
Implicit differentiation with second-order derivatives and benchmarks in finite-element-based differentiable physics
cs.CETianju Xue
Differentiable programming is revolutionizing computational science by enabling automatic differentiation (AD) of numerical simulations. While first-order gradients are well-established, second-order derivatives (Hessians) for implicit functions in finite-element-based differentiable physics remain underexplored. This work bridges this gap by deriving and im
An improved guess for the variational calculation of charge-transfer excitations in large systems
physics.chem-phNicola Bogo, Zeyi Zhang, Martin Head-Gordon, Christopher J. Stein
Charge-transfer excited states are highly relevant for applications in molecular electronics. However, the accurate calculation of these states in large systems is challenging since wave function methods are prohibitively expensive, time-dependent density functional theory with typical functionals is not precise, and the complicated topology of the electroni
Use as Many Surrogates as You Want: Selective Ensemble Attack to Unleash Transferability without Sacrificing Resource Efficiency
cs.CVBo Yang, Hengwei Zhang, Jindong Wang, Yuchen Ren
In surrogate ensemble attacks, using more surrogate models yields higher transferability but lower resource efficiency. This practical trade-off between transferability and efficiency has largely limited existing attacks despite many pre-trained models are easily accessible online. In this paper, we argue that such a trade-off is caused by an unnecessary com
Seonghyeon Moon, Young Sul Cho
This work extends the thermodynamic analysis of random bond percolation to explosive and hybrid percolation models. We show that this thermodynamic analysis is well applicable to both explosive and hybrid percolation models by using the critical exponents $\alpha$ and $\delta$ obtained from scaling relations with previously measured values of $\beta$ and $\g
Jung Hoon Lee, Sujith Vijayan
Deep learning (DL) can automatically construct intelligent agents, deep neural networks (alternatively, DL models), that can outperform humans in certain tasks. However, the operating principles of DL remain poorly understood, making its decisions incomprehensible. As a result, it poses a great risk to deploy DL in high-stakes domains in which mistakes or er
Yue Huang, Tianle Hu, Yu Chen, Zi'ang Li
Single image reflection separation aims to separate the transmission and reflection layers from a mixed image. Existing methods typically combine general priors from pre-trained models with task-specific priors such as text prompts and reflection detection. However, the transmission prior, as the most direct task-specific prior for the target transmission la
GDPRShield: AI-Powered GDPR Support for Software Developers in Small and Medium-Sized Enterprises
cs.CRTharaka Wijesundara, Mathew Warren, Nalin Arachchilage
With the rapid increase in privacy violations in modern software development, regulatory frameworks such as the General Data Protection Regulation (GDPR) have been established to enforce strict data protection practices. However, insufficient privacy awareness among SME software developers contributes to failure in GDPR compliance. For instance, a developer
Reconstruction of the occupied and unoccupied electronic states driven by quantum charge fluctuations in electron doped cuprate superconductors
cond-mat.str-elHiroshi Yamaguchi, Yudai Miyai, Yuki. Tsubota, Masashi Atira
The origin of electron-boson interactions is central to understanding high-$T_c$ superconductivity in cuprates. While phonons and magnetic fluctuations are widely considered as candidates for mediating electron pairing, the role of charge fluctuations -- one of the fundamental electronic degrees of freedom -- remains unclear. Here, we investigate the electro
ChromFound: Towards A Universal Foundation Model for Single-Cell Chromatin Accessibility Data
q-bio.GNYifeng Jiao, Yuchen Liu, Yu Zhang, Xin Guo
The advent of single-cell Assay for Transposase-Accessible Chromatin using sequencing (scATAC-seq) offers an innovative perspective for deciphering regulatory mechanisms by assembling a vast repository of single-cell chromatin accessibility data. While foundation models have achieved significant success in single-cell transcriptomics, there is currently no f
Nan Gao, Xue-Song Lu, Pu Zhang
In this article we try to recall Claus Michael Ringel's works on the Gorenstein-projective modules. This will involve but not limited to his fundamental contributions, such as in, the solution to the independence problem of totally reflexivity conditions; the technique of $\mho$-quivers; a fast algorithm to obtain the Gorenstein-projective modules over the N
Revealing the Deceptiveness of Knowledge Editing: A Mechanistic Analysis of Superficial Editing
cs.CLJiakuan Xie, Pengfei Cao, Yubo Chen, Kang Liu
Knowledge editing, which aims to update the knowledge encoded in language models, can be deceptive. Despite the fact that many existing knowledge editing algorithms achieve near-perfect performance on conventional metrics, the models edited by them are still prone to generating original knowledge. This paper introduces the concept of "superficial editing" to
MVPainter: Accurate and Detailed 3D Texture Generation via Multi-View Diffusion with Geometric Control
cs.CVMingqi Shao, Feng Xiong, Zhaoxu Sun, Mu Xu
Recently, significant advances have been made in 3D object generation. Building upon the generated geometry, current pipelines typically employ image diffusion models to generate multi-view RGB images, followed by UV texture reconstruction through texture baking. While 3D geometry generation has improved significantly, supported by multiple open-source frame
Jiazhu Li, Jian Kuang, Xiaoji Niu
To overcome the limitation of existing indoor odometry technologies which often cannot simultaneously meet requirements for accuracy cost-effectiveness, and robustness-this paper proposes a novel magnetometer array-aided inertial odometry approach, MSCEKF-MIO (Multi-State Constraint Extended Kalman Filter-based Magnetic-Inertial Odometry). We construct a mag
Alfredo Deaño, Kenneth T-R McLaughlin, Leslie Molag, Nick Simm
We carry out the asymptotic analysis as $n \to \infty$ of a class of orthogonal polynomials $p_{n}(z)$ of degree $n$, defined with respect to the planar measure \begin{equation*} d\mu(z) = (1-|z|^{2})^{\alpha-1}|z-x|^{\gamma}\mathbf{1}_{|z| < 1}d^{2}z, \end{equation*} where $d^{2}z$ is the two dimensional area measure, $\alpha$ is a parameter that can grow w
Yunseok Jang, Yeda Song, Sungryull Sohn, Lajanugen Logeswaran
Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have sparked significant interest in developing GUI visual agents. We introduce MONDAY (Mobile OS Navigation Task Dataset for Agents from YouTube), a large-scale dataset of 313K annotated frames from 20K instructional videos capturing diverse real-world mobile OS navigation
Li Lin
The 3D human pose is vital for modern computer vision and computer graphics, and its prediction has drawn attention in recent years. 3D human pose prediction aims at forecasting a human's future motion from the previous sequence. Ignoring that the arbitrariness of human motion sequences has a firm origin in transition in both temporal and spatial axes limits
Xiangpeng Tian, Xiangyu Liao, Xiao Liu, Meng Li
All-in-one image restoration aims to recover clear images from various degradation types and levels with a unified model. Nonetheless, the significant variations among degradation types present challenges for training a universal model, often resulting in task interference, where the gradient update directions of different tasks may diverge due to shared par
Yuchang Sun, Yanxi Chen, Yaliang Li, Bolin Ding
Augmenting large language models (LLMs) with auxiliary tokens has emerged as a promising strategy for enhancing model performance. In this work, we introduce a lightweight method termed latent tokens; these are dummy tokens that may be non-interpretable in natural language but steer the autoregressive decoding process of a Transformer-based LLM via the atten
Improving Generative Inverse Design of Rectangular Patch Antennas with Test Time Optimization
eess.SPBeck LaBash, Shahriar Khushrushahi, Fabian Ruehle
We propose a two-stage deep learning framework for the inverse design of rectangular patch antennas. Our approach leverages generative modeling to learn a latent representation of antenna frequency response curves and conditions a subsequent generative model on these responses to produce feasible antenna geometries. We further demonstrate that leveraging sea
Wanfu Gao, Zengyao Man, Hanlin Pan, Kunpeng Liu
Feature generation involves creating new features from raw data to capture complex relationships among the original features, improving model robustness and machine learning performance. Current methods using reinforcement learning for feature generation have made feature exploration more flexible and efficient. However, several challenges remain: first, dur
Efficient Heuristics Generation for Solving Combinatorial Optimization Problems Using Large Language Models
cs.NEXuan Wu, Di Wang, Chunguo Wu, Lijie Wen
Recent studies exploited Large Language Models (LLMs) to autonomously generate heuristics for solving Combinatorial Optimization Problems (COPs), by prompting LLMs to first provide search directions and then derive heuristics accordingly. However, the absence of task-specific knowledge in prompts often leads LLMs to provide unspecific search directions, obst
scSiameseClu: A Siamese Clustering Framework for Interpreting single-cell RNA Sequencing Data
q-bio.GNPing Xu, Zhiyuan Ning, Pengjiang Li, Wenhao Liu
Single-cell RNA sequencing (scRNA-seq) reveals cell heterogeneity, with cell clustering playing a key role in identifying cell types and marker genes. Recent advances, especially graph neural networks (GNNs)-based methods, have significantly improved clustering performance. However, the analysis of scRNA-seq data remains challenging due to noise, sparsity, a
Ali Naseh, Harsh Chaudhari, Jaechul Roh, Mingshi Wu
DeepSeek recently released R1, a high-performing large language model (LLM) optimized for reasoning tasks. Despite its efficient training pipeline, R1 achieves competitive performance, even surpassing leading reasoning models like OpenAI's o1 on several benchmarks. However, emerging reports suggest that R1 refuses to answer certain prompts related to politic
EndoForce: Development of an Intuitive Axial Force Measurement Device for Endoscopic Robotic Systems
cs.ROHansoul Kim, Dong-Ho Lee, Dukyoo Kong, Dong-Soo Kwon
Robotic endoscopic systems provide intuitive control and eliminate radiation exposure, making them a promising alternative to conventional methods. However, the lack of axial force measurement from the robot remains a major challenge, as it can lead to excessive colonic elongation, perforation, or ureteral complications. Although various methods have been pr
Lightweight and Effective Preference Construction in PIBT for Large-Scale Multi-Agent Pathfinding
cs.MAKeisuke Okumura, Hiroki Nagai
PIBT is a computationally lightweight algorithm that can be applied to a variety of multi-agent pathfinding (MAPF) problems, generating the next collision-free locations of agents given another. Because of its simplicity and scalability, it is becoming a popular underlying scheme for recent large-scale MAPF methods involving several hundreds or thousands of
Tongrui Li, Zhanfeng Liu, Peng Li, Yuzhe Wang
The interplay between magnetism and electronic band structure is a central theme in condensed matter physics. CeSb, with its complex devil's staircase antiferromagnetic transition, offers a unique opportunity to explore this interplay. Using angle-resolved photoemission spectroscopy (ARPES), we investigate the electronic structure evolution across the devil'
Keqi Deng, Philip C. Woodland
While Transformer self-attention offers strong parallelism, the Key-Value (KV) cache grows linearly with sequence length and becomes a bottleneck for inference efficiency. Multi-head latent attention was recently developed to compress the KV cache into a low-rank latent space. This paper proposes Multi-head Temporal Latent Attention (MTLA), which further red
João Eduardo Batista, Emil Vatai, Mohamed Wahib
Large Language Models (LLMs) are increasingly applied in various science domains, yet their broader adoption remains constrained by a critical challenge: the lack of trustworthy, verifiable outputs. Current LLMs often generate answers without reliable source attribution, or worse, with incorrect attributions, posing a barrier to their use in scientific and h
Forrest Mozer, Oleksiy Agapitov, Stuart Bale, John Bonnell
On November 6, 2024, the Parker Solar Probe flew past Venus to make the first accurate electric field measurement in the nightside Venusian magnetosphere. To achieve this result, the electric field antennas were current biased in a way never before experienced by an electric field detector. This biasing requirement, that the positive bias current in the Venu