October 2025 arXiv papers — page 71
Showing 7,001–7,100 of 25,213 papers
Sayantika Bhowal, Arnab Bose
Recently, spin splitting of non-relativistic origin in compensated antiferromagnets has drawn growing attention in condensed matter research. Although many materials, now known to exhibit such spin splitting, have been studied for decades, their manifestation along non-high-symmetry momentum directions initially hindered their recognition. In recent years, s
Hassan Firouzjahi
We revisit our earlier work and investigate the bound state perturbations in the interior of the Schwarzschild black hole. The bound sates are defined as the perturbations in the interior of the black hole with an imaginary spectrum which are regular at the center of black hole while their time-dependent profile falls off exponentially on the event horizon.
Bentley DeVilling
Large language models are often described as capable of reflective reasoning, yet recursive self-evaluation without external feedback frequently yields reformulation rather than progress. We test this prediction in a cross-provider study of 144 reasoning sequences across three models (OpenAI GPT-4o-mini, Anthropic Claude 3 Haiku, and Google Gemini 2.0 Flash)
Exploring Generative Process Reward Modeling for Semi-Structured Data: A Case Study of Table Question Answering
cs.CLLei Tang, Wei Zhou, Mohsen Mesgar
Process reward models (PRMs) enhance complex reasoning in large language models (LLMs) by evaluating candidate solutions step-by-step and selecting answers based on aggregated step scores. While effective in domains such as mathematics, their applicability to tasks involving semi-structured data, like table question answering (TQA), remains unexplored. TQA p
Palle Jeppesen, Bjarne Tromborg
A tutorial on opto-electronic clock regeneration at very high bit rates beyond reach with purely electronic solutions is given. Emphasis is placed on sum frequency generation in a nonlinear material such as LiNbO3. We first provide a basic introduction to CR (clock recovery) and a PLL (phase-locked loop); two examples are considered, an input signal frequenc
Jan Buchmann, Iryna Gurevych
Citations from LLM-based RAG systems are supposed to simplify response verification. However, this goal is undermined in cases of citation failure, where a model generates a helpful response, but fails to generate citations to complete evidence. In contrast to previous work, we propose to disentangle this from response failure, where the response itself is f
InvDec: Inverted Decoder for Multivariate Time Series Forecasting with Separated Temporal and Variate Modeling
cs.LGYuhang Wang
Multivariate time series forecasting requires simultaneously modeling temporal patterns and cross-variate dependencies. Channel-independent methods such as PatchTST excel at temporal modeling but ignore variable correlations, while pure variate-attention approaches such as iTransformer sacrifice temporal encoding. We proposeInvDec (Inverted Decoder), a hybri
Excluding a Line Minor via Design Matrices and Column Number Bounds for the Circuit Imbalance Measure
cs.DMDaniel Dadush, Friedrich Eisenbrand, Rom Pinchasi, Thomas Rothvoss
For a real matrix $A \in \mathbb{R}^{d \times n}$ with non-collinear columns, we show that $n \leq O(d^4 \kappa_A)$ where $\kappa_A$ is the \emph{circuit imbalance measure} of $A$. The circuit imbalance measure $\kappa$ is a real analogue of $\Delta$-modularity for integer matrices, satisfying $\kappa_A \leq \Delta_A$ for integer $A$. The circuit imbalance m
Privacy Protection of Automotive Location Data Based on Format-Preserving Encryption of Geographical Coordinates
cs.CRHaojie Ji, Long Jin, Haowen Li, Chongshi Xin
There are increasing risks of privacy disclosure when sharing the automotive location data in particular functions such as route navigation, driving monitoring and vehicle scheduling. These risks could lead to the attacks including user behavior recognition, sensitive location inference and trajectory reconstruction. In order to mitigate the data security ri
DB-FGA-Net: Dual Backbone Frequency Gated Attention Network for Multi-Class Brain Tumor Classification with Grad-CAM Interpretability
cs.LGSaraf Anzum Shreya, MD. Abu Ismail Siddique, Sharaf Tasnim
Brain tumors are a challenging problem in neuro-oncology, where early and precise diagnosis is important for successful treatment. Deep learning-based brain tumor classification methods often rely on heavy data augmentation which can limit generalization and trust in clinical applications. In this paper, we propose a double-backbone network integrating VGG16
Nonlinear stability of a composite wave to the Cauchy problem of 1-D full compressible Navier-Stokes-Allen-Cahn system
math.APDan Lei, Zhengzheng Chen
The compressible Navier-Stokes-Allen-Cahn system models the motion of a mixture of two macroscopically immiscible viscous compressible fluids. In this paper, we are concerned with the large time behavior of solutions to the Cauchy problem of the one-dimensional full compressible Navier-Stokes-Allen-Cahn system. If the Riemann problem of the corresponding Eul
Xiao Song, John Heidemann
Routing is central to networking performance, including: (1) latency in anycast services and websites served from multiple locations,(2) networking expenses and throughput in multi-homed enterprises, (3) the ability to keep traffic domestic when considering data sovereignty. However, understanding and managing how routing affects these services is challengin
Wenqi Jiang
Retrieval-augmented generation (RAG) has emerged as one of the most prominent applications of vector databases. By integrating documents retrieved from a database into the prompt of a large language model (LLM), RAG enables more reliable and informative content generation. While there has been extensive research on vector databases, many open research proble
Yang Qiu, Yixiong Zou, Jun Wang, Wei Liu
Out-of-distribution generalization under distributional shifts remains a critical challenge for graph neural networks. Existing methods generally adopt the Invariant Risk Minimization (IRM) framework, requiring costly environment annotations or heuristically generated synthetic splits. To circumvent these limitations, in this work, we aim to develop an IRM-f
Probability model of edge-fault tolerance for regular graphs with respect to edge connectivity
math.COHuanshen Jia, Jianguo Qian
We consider the probability model of edge-fault tolerance of a network in the sense of connectivity with link faults. Using graph-theoretical notation, we define the edge-fault (EF) and Menger-type edge-fault (MEF) tolerances of a graph as the probabilities that the graph is connected and strongly Menger edge-connected when each edge has a certain failure pr
Moving or Predicting? RoleAware-MAPP: A Role-Aware Transformer Framework for Movable Antenna Position Prediction to Secure Wireless Communications
cs.ITWenxu Wang, Xiaowu Liu, Wei Gong, Yujia Zhao
Movable antenna (MA) technology provides a promising avenue for actively shaping wireless channels through dynamic antenna positioning, thereby enabling electromagnetic radiation reconstruction to enhance physical layer security (PLS). However, its practical deployment is hindered by two major challenges: the high computational complexity of real time optimi
Callum Sharrock, Lukas Petersson, Hanna Petersson, Axel Backlund
We present Butter-Bench, a benchmark evaluating large language model (LLM) controlled robots for practical intelligence, defined as the ability to navigate the messiness of the physical world. Current state-of-the-art robotic systems use a hierarchical architecture with LLMs in charge of high-level reasoning, and a Vision Language Action (VLA) model for low-
Vincent Moulton, Andreas Spillner
In 1989 Erd\H{o}s and Sz\'ekely showed that there is a bijection between (i) the set of rooted trees with $n+1$ vertices whose leaves are bijectively labeled with the elements of $[\ell]=\{1,2,\dots,\ell\}$ for some $\ell \leq n$, and (ii) the set of partitions of $[n]=\{1,2,\dots,n\}$. They established this via a labeling algorithm based on the anti-lexicog
LinFeng Li, Jian Zhao, Zepeng Yang, Yuhang Song
We present a winning solution to RoboSense 2025 Track 4: Cross-Modal Drone Navigation. The task retrieves the most relevant geo-referenced image from a large multi-platform corpus (satellite/drone/ground) given a natural-language query. Two obstacles are severe inter-platform heterogeneity and a domain gap between generic training descriptions and platform-s
On an Analytical Criterion for Detecting Intermittent Turbulent Behaviour of Solutions of Partial Differential Equations
math.APMichele V Bartuccelli, Guido Gentile
A main question in the study of partial differential equations is the following: how do we understand the nature of the solutions and, in particular, how do we determine if a given solution shows turbulent or non-turbulent behaviour? Being able to answer such a question would be a major advance in the comprehension of the nature of turbulence. In this paper
Jinhong Zhao, Bin Guo
We study a one-dimensional nonlocal degenerate fourth-order parabolic equation with inhomogeneous forces relevant to hydraulic fracture modeling. Employing a regularization scheme, modified energy/entropy methods, and novel differential inequality techniques, we establish global existence and long-time behavior results for weak solutions under both time-and
Yingxi Li, Ellen Vitercik, Mingwei Yang
In the online metric matching problem, $n$ servers and $n$ requests lie in a metric space. Servers are available upfront, and requests arrive sequentially. An arriving request must be matched immediately and irrevocably to an available server, incurring a cost equal to their distance. The goal is to minimize the total matching cost. We study this problem in
Sauptik Dhar, Naveen Ramakrishnan, Michelle Munson
Large Vision Language models have seen huge application in several sports use-cases recently. Most of these works have been targeted towards a limited subset of popular sports like soccer, cricket, basketball etc; focusing on generative tasks like visual question answering, highlight generation. This work analyzes the applicability of the modern video founda
Liangyu Chen, Hanzhang Zhou, Chenglin Cai, Jianan Zhang
GUI grounding, which maps natural-language instructions to actionable UI elements, is a core capability of GUI agents. Prior works largely treats instructions as a static proxy for user intent, overlooking the impact of instruction diversity and quality on grounding performance. Through a careful investigation of existing grounding datasets, we find a 23.3%
DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
cs.CVJiayi Zou, Chaofan Chen, Bing-Kun Bao, Changsheng Xu
Egocentric Video Question Answering (Egocentric VideoQA) plays an important role in egocentric video understanding, which refers to answering questions based on first-person videos. Although existing methods have made progress through the paradigm of pre-training and fine-tuning, they ignore the unique challenges posed by the first-person perspective, such a
Haodong Yang, Zhongling Huang, Shaojie Guo, Zhe Zhang
Deep learning models for complex-valued Synthetic Aperture Radar (CV-SAR) image recognition are fundamentally constrained by a representation trilemma under data-limited and domain-shift scenarios: the concurrent, yet conflicting, optimization of generalization, interpretability, and efficiency. Our work is motivated by the premise that the rich electromagne
Jiayi Zou, Gengyun Jia, Bing-Kun Bao
Visual Commonsense Reasoning (VCR) refers to answering questions and providing explanations based on images. While existing methods achieve high prediction accuracy, they often overlook bias in datasets and lack debiasing strategies. In this paper, our analysis reveals co-occurrence and statistical biases in both textual and visual data. We introduce the VCR
Beiya Dai, Yuliang Liu, Daozheng Xue, Yunchong Song
We propose ContextLM, a framework that implicitly learns multi-token prediction by augmenting standard pretraining with an intrinsic next-context prediction objective. ContextLM builds a language model on top of context embeddings that span multiple tokens, enabling better next-token prediction by predicting the next context. Our model is fully compatible wi
Penghao Wang, Yuhao Zhou, Mengxuan Wu, Ziheng Qin
As large language models (LLMs) advance, the ultimate vision for their role in science is emerging: we could build an AI collaborator to effectively assist human beings throughout the entire scientific research process. We refer to this envisioned system as ResearchGPT. Given that scientific research progresses through multiple interdependent phases, achievi
Guangyu Dai, Siliang Tang, Yueting Zhuang
In recent years, Pretrained Large Models(PLMs) researchers proposed large-small model collaboration frameworks, leveraged easily trainable small models to assist large models, aim to(1) significantly reduce computational resource consumption while maintaining comparable accuracy, and (2) enhance large model performance in specialized domain tasks. However, t
A Location-Aware Hybrid Deep Learning Framework for Dynamic Near-Far Field Channel Estimation in Low-Altitude UAV Communications
cs.ITWenli Yuan, Kan Yu, Xiaowu Liu, Kaixuan Li
In low altitude UAV communications, accurate channel estimation remains challenging due to the dynamic nature of air to ground links, exacerbated by high node mobility and the use of large scale antenna arrays, which introduce hybrid near and far field propagation conditions. While conventional estimation methods rely on far field assumptions, they fail to c
Wonil Kim, Hyeongseok Wi, Seungsoon Park, Taejun Kim
Generative AI is reshaping music creation, but its rapid growth exposes structural gaps in attribution, rights management, and economic models. Unlike past media shifts, from live performance to recordings, downloads, and streaming, AI transforms the entire lifecycle of music, collapsing boundaries between creation, distribution, and monetization. However, e
Yunzhi Liu, Haokai Tan, Rushi Kanjaria, Lihuan Li
Human mobility forecasting is crucial for disaster relief, city planning, and public health. However, existing models either only model location sequences or include time information merely as auxiliary input, thereby failing to leverage the rich semantic context provided by points of interest (POIs). To address this, we enrich a BERT-based mobility model wi
Kangda Zhi, Tianyu Yang, Songyan Xue, Giuseppe Caire
This paper investigates the design of channel estimation and 3D localization algorithms in a challenging scenario, where a sub-connected planar extremely large-scale multiple-input multiple-output (XL-MIMO) communicates with multi-antenna users. In the near field, the uplink MIMO channel is of full column rank and therefore can not be estimated effectively b
Qitai Tan, Yiyun Chen, Mo Li, Ruiwen Gu
Recent advances in deep learning have driven rapid progress in time series forecasting, yet many state-of-the-art models continue to struggle with robust performance in real-world applications, even when they achieve strong results on standard benchmark datasets. This persistent gap can be attributed to the black-box nature of deep learning architectures and
Tristan Cinquin, Geoff Pleiss, Agustinus Kristiadi
While chain-of-thought prompting with Best-of-N (BoN) selection has become popular for mathematical reasoning in large language models (LLMs), its linear structure fails to capture the branching and exploratory nature of complex problem-solving. In this work, we propose an adaptive algorithm to maximize process reward model (PRM) scores over the intractable
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
cs.LGUdit Saxena
Topological features capture global geometric structure in imaging data, but practical adoption in deep learning requires both computational efficiency and differentiability. We present optimized GPU kernels for the Euler Characteristic Curve (ECC) computation achieving 16-2000\"O speedups over prior GPU implementations on synthetic grids, and introduce a di
Ziqian Zhong, Aditi Raghunathan, Nicholas Carlini
The tendency to find and exploit "shortcuts" to complete tasks poses significant risks for reliable assessment and deployment of large language models (LLMs). For example, an LLM agent with access to unit tests may delete failing tests rather than fix the underlying bug. Such behavior undermines both the validity of benchmark results and the reliability of r
In-DRAM True Random Number Generation Using Simultaneous Multiple-Row Activation: An Experimental Study of Real DRAM Chips
cs.ARIsmail Emir Yuksel, Ataberk Olgun, F. Nisa Bostanci, Oguzhan Canpolat
In this work, we experimentally demonstrate that it is possible to generate true random numbers at high throughput and low latency in commercial off-the-shelf (COTS) DRAM chips by leveraging simultaneous multiple-row activation (SiMRA) via an extensive characterization of 96 DDR4 DRAM chips. We rigorously analyze SiMRA's true random generation potential in t
A Principle-based Framework for the Development and Evaluation of Large Language Models for Health and Wellness
cs.HCBrent Winslow, Jacqueline Shreibati, Javier Perez, Hao-Wei Su
The incorporation of generative artificial intelligence into personal health applications presents a transformative opportunity for personalized, data-driven health and fitness guidance, yet also poses challenges related to user safety, model accuracy, and personal privacy. To address these challenges, a novel, principle-based framework was developed and val
Guangyu Dai, Dong Chen, Siliang Tang, Yueting Zhuang
Video anomaly detection (VAD) is a challenging task that detects anomalous frames in continuous surveillance videos. Most previous work utilizes the spatio-temporal correlation of visual features to distinguish whether there are abnormalities in video snippets. Recently, some works attempt to introduce multi-modal information, like text feature, to enhance t
Saraf Anzum Shreya, MD. Abu Ismail Siddique, Sharaf Tasnim
Technologies like smartphones have become an essential in our daily lives. It has made accessible to everyone including visually impaired individuals. With the use of smartphone cameras, image capturing and processing have become more convenient. With the use of smartphones and machine learning, the life of visually impaired can be made a little easier. Dail
Mahtab Movaheddrad, Laurence Palmer, C. -C. Jay Kuo
Image dehazing is a restoration task that aims to recover a clear image from a single hazy input. Traditional approaches rely on statistical priors and the physics-based atmospheric scattering model to reconstruct the haze-free image. While recent state-of-the-art methods are predominantly based on deep learning architectures, these models often involve high
A Survey of OTFS-Based Index Modulation Techniques: Challenges, Benefits, and Future Directions for 6G and Beyond
eess.SPBurak Ahmet Ozden, Erdogan Aydin, Emir Aslandogan, Haci Ilhan
Orthogonal time frequency space (OTFS) is a two-dimensional modulation technique that uses the delay-Doppler (DD) domain and is a candidate for providing robust, high-capacity wireless communications for envisioned 6G and beyond networks. The OTFS technique maps data to the DD domain instead of the traditional time-frequency domain, enabling it to fully util
Thomas Rupf, Marco Bagatella, Marin Vlastelica, Andreas Krause
Behavior Foundation Models (BFMs) are capable of retrieving high-performing policy for any reward function specified directly at test-time, commonly referred to as zero-shot reinforcement learning (RL). While this is a very efficient process in terms of compute, it can be less so in terms of data: as a standard assumption, BFMs require computing rewards over
Tahir Boudjeriou, Prosenjit Roy
In this paper, we are concerned with the asymptotic behavior of weak solutions to certain elliptic and parabolic problems involving the fractional $p$-Laplacian in cylindrical domains that become unbounded in one direction. The nonlocal nature of the operator describing the equations creates several technical difficulties in treating problems of this type. T
Natalia S. Salakhova, Sergey A. Dyakov, Ilia M. Fradkin, Nikolay A. Gippius
We investigate the Casimir effect in a system of two twisted photonic gratings made of uniaxially anisotropic materials. Two distinct configuretions are explored: a stack of symmetric gratings and a stack of in-plane chiral gratings, with the latter realized by choosing specific orientaton of anisotropy axis relative to stripes. We apply the reflection-matri
Mert Bulent Sariyildiz, Philippe Weinzaepfel, Guillaume Bono, Gianluca Monaci
One key aspect of spatially aware robots is the ability to "find their bearings", ie. to correctly situate themselves in previously seen spaces. In this work, we focus on this particular scenario of continuous robotics operations, where information observed before an actual episode start is exploited to optimize efficiency. We introduce a new model, Kinaema,
Changping Meng, Hongyi Ling, Jianling Wang, Yifan Liu
Large Language Models (LLMs) empower recommendation systems through their advanced reasoning and planning capabilities. However, the dynamic nature of user interests and content poses a significant challenge: While initial fine-tuning aligns LLMs with domain knowledge and user preferences, it fails to capture such real-time changes, necessitating robust upda
Bowen Gang, Hongmei Lin, Tiejun Tong
Tukey's boxplot is a foundational tool for exploratory data analysis, but its classic outlier-flagging rule does not account for the sample size, and subsequent modifications have often been presented as separate, heuristic adjustments. In this paper, we propose a unifying framework that recasts the boxplot and its variants as graphical implementations of mu
Bita Banihashemi, Megh Patel, Yves Lespérance
Generating an abstraction of a dynamic domain that aligns with a given purpose remains a significant challenge given that the choice of such an abstraction can impact an agent's ability to plan, reason, and provide explanations effectively. We model the agent's concrete behaviors in PDDL and investigate the use of in-context learning with large language mode
Guowei Zhong, Junjie Li, Huaiyu Zhu, Ruohong Huan
In recent years, Multimodal Emotion Recognition (MER) has made substantial progress. Nevertheless, most existing approaches neglect the semantic inconsistencies that may arise across modalities, such as conflicting emotional cues between text and visual inputs. Besides, current methods are often dominated by the text modality due to its strong representation
Yogesh Simmhan, Varad Kulkarni
This article presents early findings from designing, deploying and evaluating an AI-based educational agent deployed as the primary instructor in a graduate-level Cloud Computing course at IISc. We detail the design of a Large Language Model (LLM)-driven Instructor Agent, and introduce a pedagogical framework that integrates the Instructor Agent into the cou
Kazufumi Ito, Tiancheng Xue
In this paper we consider the neural network optimization. We develop Anderson-type acceleration method for the stochastic gradient decent method and it improves the network permanence very much. We demonstrate the applicability of the method for Deep Neural Network (DNN) and Convolution Neural Network (CNN).
Zan Li, Rui Fan
Financial time series forecasting faces a fundamental challenge: predicting optimal asset allocations requires understanding regime-dependent correlation structures that transform during crisis periods. Existing graph-based spatio-temporal learning approaches rely on predetermined graph topologies--correlation thresholds, sector classifications--that fail to
Individualized Cognitive Simulation in Large Language Models: Evaluating Different Cognitive Representation Methods
cs.AITianyi Zhang, Xiaolin Zhou, Yunzhe Wang, Erik Cambria
Individualized cognitive simulation (ICS) aims to build computational models that approximate the thought processes of specific individuals. While large language models (LLMs) convincingly mimic surface-level human behavior such as role-play, their ability to simulate deeper individualized cognitive processes remains poorly understood. To address this gap, w
Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards
cs.LGJiajun Fan, Roger Ren, Jingyuan Li, Rahul Pandey
The role of reasoning in Audio Large Language Models remains widely underexplored, as introducing a reasoning process often degrades rather than improves performance during inference, a phenomenon we term test-time inverse scaling, where longer reasoning chains yield progressively worse results. We demonstrate that this stems not from fundamental limitations
Jie Wen, Benjian Lv
A $k$-partition of an $n$-set $X$ is a collection of $k$ pairwise disjoint non-empty subsets whose union is $X$. A family of $k$-partitions of $X$ is called $t$-intersecting if any two of its members share at least $t$ blocks. A $t$-intersecting family is trivial if every $k$-partition in it contains $t$ fixed blocks, and is non-trivial otherwise. In this pa
Zhiqin Yang, Yonggang Zhang, Chenxin Li, Yiu-ming Cheung
Federated Learning (FL) confronts a significant challenge known as data heterogeneity, which impairs model performance and convergence. Existing methods have made notable progress in addressing this issue. However, improving performance in certain heterogeneity scenarios remains an overlooked question: \textit{How robust are these methods to deploy under div
Yicao Wang
This paper is a continuation of our previous work \cite{wang2024complex}. It mainly deals with entire operators $T$ with deficiency index 1 \emph{systematically} from the complex-geometric viewpoint proposed in \cite{wang2024complex}. We pay special attention to the characteristic line bundle $F$ of $T$. We investigate its curvature in detail and demonstrate
Sarah E. Kay, Brian N. Wenny
Underflight maneuvers provide a unique opportunity to harmonize calibration of on-orbit sensors. Due to their similar sensor technologies, their near-identical transmission profiles, orbital properties and platform operations, the underflight data of Landsat 8 and 9 instruments stand out as a qualifier to test proposed metrics, methods, and the extent over w
Seeing the Unseen: Mask-Driven Positional Encoding and Strip-Convolution Context Modeling for Cross-View Object Geo-Localization
cs.CVShuhan Hu, Yiru Li, Yuanyuan Li, Yingying Zhu
Cross-view object geo-localization enables high-precision object localization through cross-view matching, with critical applications in autonomous driving, urban management, and disaster response. However, existing methods rely on keypoint-based positional encoding, which captures only 2D coordinates while neglecting object shape information, resulting in s
Yan Guo, Heng Lyu, Chunling Ding, Chenzhi Yuan
Fractional-order vortex beams possess fractional orbital angular momentum (FOAM) modes, which theoretically have the potential to increase transmission capacity infinitely. Therefore, they have significant application prospects in the fields of measurement, optical communication and micro-particle manipulation. However, when fractional-order vortex beams pro
Minseok Kang, Minhyeok Lee, Minjung Kim, Donghyeong Kim
Video Temporal Grounding (VTG) aims to localize temporal segments in long, untrimmed videos that align with a given natural language query. This task typically comprises two subtasks: Moment Retrieval (MR) and Highlight Detection (HD). While recent advances have been progressed by powerful pretrained vision-language models such as CLIP and InternVideo2, exis
Yu Hin Chan, Hao Yang, Shiyu Shen, Xingyu Fan
Privacy-preserving machine learning (PPML) is an emerging topic to handle secure machine learning inference over sensitive data in untrusted environments. Fully homomorphic encryption (FHE) enables computation directly on encrypted data on the server side, making it a promising approach for PPML. However, it introduces significant communication and computati
Stephan Rabanser, Nicolas Papernot
Selective classifiers improve model reliability by abstaining on inputs the model deems uncertain. However, few practical approaches achieve the gold-standard performance of a perfect-ordering oracle that accepts examples exactly in order of correctness. Our work formalizes this shortfall as the selective-classification gap and present the first finite-sampl
New Second-Order Achievability Bounds for Coding with Side Information via Type Deviation Convergence
cs.ITXiang Li, Cheuk Ting Li
We propose a framework for second-order achievability, called type deviation convergence, that is generally applicable to settings in network information theory, and is especially suitable for lossy source coding and channel coding with cost. We give a second-order achievability bound for lossy source coding with side information at the decoder (Wyner-Ziv pr
Filippo Cenacchi, Deborah Richards, Longbing Cao
Depression and post traumatic stress disorder (PTSD) often co-occur with connected symptoms, complicating automated assessment, which is often binary and disorder specific. Clinically useful diagnosis needs severity aware cross disorder estimates and decision support explanations. Our unified tri modal affective severity framework synchronizes and fuses inte
Runsong Zhu, Ka-Hei Hui, Zhengzhe Liu, Qianyi Wu
Open-vocabulary 3D segmentation is a fundamental yet challenging task, requiring a mutual understanding of both segmentation and language. However, existing Gaussian-splatting-based methods rely either on a single 3D language field, leading to inferior segmentation, or on pre-computed class-agnostic segmentations, suffering from error accumulation. To addres
Electric field induced Berry curvature dipole and non-linear anomalous Hall effects in higher wave symmetric unconventional magnets
cond-mat.mes-hallSrimayi Korrapati, Snehasish Nandy, Sumanta Tewari
We investigate the second-order anomalous Hall response in two-dimensional higher-wave symmetric magnets, including the recently discovered class of collinear magnets known as altermagnets, when subjected to a symmetry-breaking external electric field. In these systems, the first- and second-order anomalous Hall responses mediated by the first- and second-or
Teng Jiek See, Daokun Zhang, Mario Boley, David K. Chalmers
Graph Neural Networks (GNNs) are the currently most effective methods for predicting molecular properties but there remains a need for more accurate models. GNN accuracy can be improved by increasing the model complexity but this also increases the computational cost and memory requirement during training and inference. In this study, we develop Layer-to-Lay
Woohyeon Byeon, Giseung Park, Jongseong Chae, Amir Leshem
In this paper, we propose a provably convergent and practical framework for multi-objective reinforcement learning with max-min criterion. From a game-theoretic perspective, we reformulate max-min multi-objective reinforcement learning as a two-player zero-sum regularized continuous game and introduce an efficient algorithm based on mirror descent. Our appro
Van Huynh, Hieu Trinh, Riley Bain
We present a collection of different types of observation systems that work as differentiators. These observer-based differentiators can produce estimates for derivatives of a given signal, even though the given signal is prone to noise.
Khai-Hoan Nguyen-Dang, Nguyen Thai Hung
Let $n\ge 2$, let $m\in\mathbb Z\setminus\{0\}$, and let $K=\mathbb Q(α)$, where $α^n=m$ and $X^n-m$ is irreducible over $\mathbb Q$. We study when the natural order $\mathbb Z[α]$ is the full ring of integers $\mathcal O_K$. For the pure family $X^n-m$, we give a short proof, using only Dedekind's index criterion, of the equivalence $\mathcal O_K = \mat
HGraphScale: Hierarchical Graph Learning for Autoscaling Microservice Applications in Container-based Cloud Computing
cs.DCZhengxin Fang, Hui Ma, Gang Chen, Rajkumar Buyya
Microservice architecture has become a dominant paradigm in application development due to its advantages of being lightweight, flexible, and resilient. Deploying microservice applications in the container-based cloud enables fine-grained elastic resource allocation. Autoscaling is an effective approach to dynamically adjust the resource provisioned to conta
NODA-MMH: Certified Learning-Aided Nonlinear Control for Magnetically-Actuated Swarm Experiment Toward On-Orbit Proof
cs.ROYuta Takahashi, Atsuki Ochi, Yoichi Tomioka, Shin-Ichiro Sakai
This study experimentally validates the principle of large-scale satellite swarm control through learning-aided magnetic field interactions generated by satellite-mounted magnetorquers. This actuation presents a promising solution for the long-term formation maintenance of multiple satellites and has primarily been demonstrated in ground-based testbeds for t
Yifan Wang, Chenchao Xu, Zhimian Wu, Huachen Rao
A range of of unusual emergent behaviors have been reported in the charge-density wave (CDW) state of the $A$V$_3$Sb$_5$ ($A=~$K, Rb, Cs) kagome metals, including a CDW formation process without soft phonons, which points to an unconventional CDW mechanism. Here, we use inelastic x-ray scattering to show that the CDW in KV$_3$Sb$_5$ forms via phonons that so
Ge Zheng, Jiaye Qian, Jiajin Tang, Sibei Yang
Large Vision-Language Models (LVLMs) have made significant progress in recent years but are also prone to hallucination issues. They exhibit more hallucinations in longer, free-form responses, often attributed to accumulated uncertainties. In this paper, we ask: Does increased hallucination result solely from length-induced errors, or is there a deeper under
Christopher S. Bird, Antonio P. L. Bo, Wei Lu, Taylor J. M. Dick
Accurate measures of musculoskeletal forces are critical for clinicians, biomechanists, and engineers, yet direct measurement is highly invasive and current estimation methods remain limited in accuracy. Here, we demonstrate the application of ultra-wideband radar to non-invasively estimate musculoskeletal forces by measuring changes in the electromagnetic p
Yago del Valle Inclan Redondo, Enrique Arriaga-Varela, Dmitry Lyamzin, Pablo Cervantes
We introduce SpLIIF to generate implicit neural representations and enable arbitrary downscaling of weather variables. We train a model from sparse weather stations and topography over Japan and evaluate in- and out-of-distribution accuracy predicting temperature and wind, comparing it to both an interpolation baseline and CorrDiff. We find the model to be u
Jinho Cha, Youngchul Kim, Jungmin Shin, Jaeyoung Cho
We develop a general optimization-theoretic framework for Bregman-Variational Learning Dynamics (BVLD), a new class of operator-based updates that unify Bayesian inference, mirror descent, and proximal learning under time-varying environments. Each update is formulated as a variational optimization problem combining a smooth convex loss f_t with a Bregman di
Bijo S. Anand, Manoj Changat, Prasanth G. Narasimha-Shenoi, Mary Shalet Thottungal Joseph
Suppose $D = (V, E)$ is a strongly connected digraph and $u, v \in V (D)$. Among the many metrics in graphs, the sum metric warrants further exploration. The sum distance $sd(u, v)$ defined as $sd(u, v) =\overrightarrow{d}(u, v)+\overrightarrow{d}(v, u)$ is a metric where $\overrightarrow{d}(u, v)$ denotes the length of the shortest directed $u - v$ path in
Insu Jeon, Minui Hong, Junhyeog Yun, Gunhee Kim
Federated Learning (FL) aims to train a global inference model from remotely distributed clients, gaining popularity due to its benefit of improving data privacy. However, traditional FL often faces challenges in practical applications, including model overfitting and divergent local models due to limited and non-IID data among clients. To address these issu
Suguru Endo, Hideaki Hakoshima, Tomohiro Shitara
Non-Markovian dynamics are typically present in the dynamics of open quantum systems. Despite the rich structure of non-Markovian dynamics, their relevance to quantum information processing (QIP) has been rarely discussed. In this work, we demonstrate that the negativity of the dynamics, a characteristic of non-Markovian dynamics, naturally arises in quantum
Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations
cs.CRDivyanshu Kumar, Shreyas Jena, Nitin Aravind Birur, Tanay Baswa
Multimodal large language models (MLLMs) have achieved remarkable progress, yet remain critically vulnerable to adversarial attacks that exploit weaknesses in cross-modal processing. We present a systematic study of multimodal jailbreaks targeting both vision-language and audio-language models, showing that even simple perceptual transformations can reliably
Zhongyi Yu, Jianqiu Wu, Zhenghao Wu, Shuhan Zhong
Temporal graph link prediction aims to predict future interactions between nodes in a graph based on their historical interactions, which are encoded in node embeddings. We observe that heterogeneity naturally appears in temporal interactions, e.g., a few node pairs can make most interaction events, and interaction events happen at varying intervals. This le
Alejandro Michel, Abhinav Arun, Bhaskarjit Sarmah, Stefano Pasquali
Portfolio managers rely on correlation-based analysis and heuristic methods that fail to capture true causal relationships driving performance. We present a hybrid framework that integrates statistical causal discovery algorithms with domain knowledge from two complementary sources: a financial knowledge graph extracted from SEC 10-K filings and large langua
Ke Xing, Yanjie Dong, Xiaoyi Fan, Runhao Zeng
Personalized federated learning (PFL) addresses a critical challenge of collaboratively training customized models for clients with heterogeneous and scarce local data. Conventional federated learning, which relies on a single consensus model, proves inadequate under such data heterogeneity. Its standard aggregation method of weighting client updates heurist
Qinyu Xu, Yuanyang Zhu, Xuefei Wu, Chunlin Chen
The ability to model interactions among agents is crucial for effective coordination and understanding their cooperation mechanisms in multi-agent reinforcement learning (MARL). However, previous efforts to model high-order interactions have been primarily hindered by the combinatorial explosion or the opaque nature of their black-box network structures. In
Jiahuan Wang, Yuxin Chen, Jun Yu, Guangming Lu
Adapting pretrained diffusion-based generative models for text-driven image editing with negligible tuning overhead has demonstrated remarkable potential. A classical adaptation paradigm, as followed by these methods, first infers the generative trajectory inversely for a given source image by image inversion, then performs image editing along the inferred t
Skyler Palatnick, Maxwell A. Millar-Blanchaer, Jingwen Zhang, Kellen Lawson
We report the discovery of a debris disk surrounding the M3 star, TWA 20, revealed by JWST coronagraphic observations using the Near-infrared Camera (NIRCam). With reference-star differential imaging (RDI), we resolve the disk in scattered light in the F200W filter at a high signal-to-noise ratio and in the F444W filter at a low signal-to-noise ratio. The di
Xuesong Wang
Rapidly increasing demand for high speed data is pushing 6G wireless networks to support larger link scales, lower latency, and higher spectral efficiency. Visible light communications (VLC) is a strong complement to radio frequency (RF) systems within 6G. The latest ITU G.9991 and IEEE 802.11bb standards are adapted from cable and RF wireless technologies f
Towards Objective Obstetric Ultrasound Assessment: Contrastive Representation Learning for Fetal Movement Detection
cs.CVTalha Ilyas, Duong Nhu, Allison Thomas, Arie Levin
Accurate fetal movement (FM) detection is essential for assessing prenatal health, as abnormal movement patterns can indicate underlying complications such as placental dysfunction or fetal distress. Traditional methods, including maternal perception and cardiotocography (CTG), suffer from subjectivity and limited accuracy. To address these challenges, we pr
Yanghao Wang, Zhen Wang, Long Chen
Recent pre-trained text-to-image flow models have enabled remarkable progress in text-based image editing. Mainstream approaches adopt a corruption-then-restoration paradigm, where the source image is first corrupted into an editable ``intermediate state'' and then restored to the target image under the prompt guidance. However, current methods construct thi
Zhenning Yang, Hui Guan, Victor Nicolet, Brandon Paulsen
Cloud infrastructure is managed through a mix of interfaces -- traditionally, cloud consoles, command-line interfaces (CLI), and SDKs are the tools of choice. Recently, Infrastructure-as-Code/IaC frameworks (e.g., Terraform) have quickly gained popularity. Unlike conventional tools, IaC~frameworks encode the infrastructure in a "source-of-truth" configuratio
Hualei Wang, Na Li, Chuke Wang, Shu Wu
Recent advances in zero-shot text-to-speech (TTS), driven by language models, diffusion models and masked generation, have achieved impressive naturalness in speech synthesis. Nevertheless, stability and fidelity remain key challenges, manifesting as mispronunciations, audible noise, and quality degradation. To address these issues, we introduce Vox-Evaluato
Assessing the Feasibility of Early Cancer Detection Using Routine Laboratory Data: An Evaluation of Machine Learning Approaches on an Imbalanced Dataset
cs.LGShumin Li
The development of accessible screening tools for early cancer detection in dogs represents a significant challenge in veterinary medicine. Routine laboratory data offer a promising, low-cost source for such tools, but their utility is hampered by the non-specificity of individual biomarkers and the severe class imbalance inherent in screening populations. T
David Pohl, Marco Cognetta, Junyoung Lee, Naoaki Okazaki
Modern language models operate on subword-tokenized text in order to make a trade-off between model size, inference speed, and vocabulary coverage. A side effect of this is that, during inference, models are evaluated by measuring the probability of only the specific tokenization produced as the output, despite there being many possible ways to represent the
Sung-Sik Lee
In this paper, we further develop a recently proposed theory of time based on wavefunction collapse in general relativity. It is based on the postulations that quantum states, which violate the momentum and Hamiltonian constraints, represent instances of time, and stochastic fluctuations of the lapse and shift generate the time evolution under which an initi
Maggie Bai, Ava Kim Cohen, Eleanor Koss, Charlie Lichtenbaum
Optimizing artificial intelligence (AI) for dynamic environments remains a fundamental challenge in machine learning research. In this paper, we examine evolutionary training methods for optimizing AI to solve the game 2048, a 2D sliding puzzle. 2048, with its mix of strategic gameplay and stochastic elements, presents an ideal playground for studying decisi