Skip to content

May 2025 arXiv papers — page 36

Showing 3,5013,600 of 24,552 papers

  1. Zijing Hu, Fengda Zhang, Kun Kuang

    The practical applications of diffusion models have been limited by the misalignment between generated images and corresponding text prompts. Recent studies have introduced direct preference optimization (DPO) to enhance the alignment of these models. However, the effectiveness of DPO is constrained by the issue of visual inconsistency, where the significant

  2. Misaki Mitsuno, Xiao Ma, Koji Hasegawa

    This paper presents the evaporation-induced freezing dynamics of pure cyclohexane droplets levitated via acoustic levitation. Acoustic levitation has attracted considerable attention across various fields owing to its potential to create lab-in-a-drop systems. While droplet evaporation is a fundamental physicochemical process in such a platform, the freezing

  3. Xiaochen Wei, Weiwei Guo, Wenxian Yu

    The substantial modality-induced variations in radiometric, texture, and structural characteristics pose significant challenges for the accurate registration of multimodal images. While supervised deep learning methods have demonstrated strong performance, they often rely on large-scale annotated datasets, limiting their practical application. Traditional un

  4. Ashim Gupta, Maitrey Mehta, Zhichao Xu, Vivek Srikumar

    Large language models (LLMs) provide detailed and impressive responses to queries in English. However, are they really consistent at responding to the same query in other languages? The popular way of evaluating for multilingual performance of LLMs requires expensive-to-collect annotated datasets. Further, evaluating for tasks like open-ended generation, whe

  5. Yuhao Hu

    For hyperbolic Monge-Amp\`ere systems, a well-known solution of the equivalence problem yields two invariant tensors, ${S}_1$ and ${S}_2$, defined on the underlying $5$-manifold, where ${S}_2=0$ characterizes systems that are Euler-Lagrange. In this article, we consider the `opposite' case, ${S}_1 = 0$, and show that the local generality of such systems is `

  6. Jihong Zhang, Xinya Liang, Anqi Deng, Nicole Bonge

    Mixed methods research integrates quantitative and qualitative data but faces challenges in aligning their distinct structures, particularly in examining measurement characteristics and individual response patterns. Advances in large language models (LLMs) offer promising solutions by generating synthetic survey responses informed by qualitative data. This s

  7. Leonid S. Taran, Anastasia E. Lebedeva, Sergey V. Streltsov

    This work focuses on the layered perovskite Sr$_2$NbO$_4$, a 4$d$ analogue of Sr$_2$VO$_4$, which remains an unsolved puzzle with a possible intriguing hidden magnetic order. Using density functional theory (DFT) calculations, we demonstrate the robust thermodynamic stability and exfoliability of Sr$_2$NbO$_4$, suggesting potential applications as a 2D mater

  8. Josef Dick, Seungchan Ko, Quoc Thong Le Gia, Kassem Mustapha

    Due to divergence instability, the accuracy of low-order conforming finite element methods for nearly incompressible elasticity equations deteriorates as the Lam\'e coefficient $\lambda\to\infty$, or equivalently as the Poisson ratio $\nu\to1/2$. This phenomenon, known as locking or non-robustness, remains not fully understood despite extensive investigation

  9. Dimpi, Hemant Kumar Singh

    The orbit spaces of free S^0-actions on the mod 2 cohomology product of three spheres, S^n x S^m x S^l, 1 <= n <= m <= l have been determined in [6]. In this paper, we extend these findings to free S^1-actions on the rational cohomology product of three spheres. This extension also builds upon the work of Dotzel et al. [7], who studied free circle actions on

  10. Hanseong Jo, Pavel Shafirin, Christopher Le, Caden Chan

    Soft electrothermal actuators are of great interest in diverse application domains for their simplicity, compliance, and ease of control. However, the very nature of thermally induced mechanical actuation sets inherent operation constraints: unidirectional motion, environmental sensitivity, and slow response times limited by passive cooling. To overcome thes

  11. Zhixing Huang, Yi Mei, Fangfang Zhang, Mengjie Zhang

    Genetic programming has undergone rapid development in recent years. However, theoretical studies of genetic programming are far behind. One of the major obstacles to theoretical studies is the challenge of developing a model to describe the relationship between fitness values and program genotypes. In this paper, we take linear genetic programming (LGP) as

  12. Zijian Zhou, Jingze Ding, Rui Zhang

    In this paper, we propose a new form of polarization reconfigurable antennas (PRAs) that can form linear, circular, and general elliptical polarizations assisted by phase shifters (PSs). With PRAs, polarforming is achieved, which enables the antenna to shape its polarization into a desired state for aligning with that of the received electromagnetic (EM) wav

  13. Bishnu Paudel, James A. Sellers, Haiyang Wang

    Recently, Alanazi, Munagi, and Saikia employed the theory of modular forms to investigate the arithmetic properties of the function $\overline{R_{\ell,\mu}}(n)$, which enumerates the overpartitions of $n$ where no part is divisible by either $\ell$ or $\mu$, for various integer pairs $(\ell, \mu)$. In this paper, we substantially extend several of their resu

  14. Ziyang Zheng, Kezhi Li, Zhengyuan Shi, Qiang Xu

    Subgraph matching in logic circuits is foundational for numerous Electronic Design Automation (EDA) applications, including datapath optimization, arithmetic verification, and hardware trojan detection. However, existing techniques rely primarily on structural graph isomorphism and thus fail to identify function-related subgraphs when synthesis transformatio

  15. Zhendong Mi, Zhenglun Kong, Geng Yuan, Shaoyi Huang

    With the rapid expansion of large language models (LLMs), the demand for memory and computational resources has grown significantly. Recent advances in LLM pruning aim to reduce the size and computational cost of these models. However, existing methods often suffer from either suboptimal pruning performance or low time efficiency during the pruning process.

  16. Kanta Takahata, Jonas Schöpf, Naoki Nishida, Takahito Aoto

    Logically constrained term rewriting is a rewriting framework that supports built-in data structures such as integers and bit vectors. Recently, constrained terms play a key role in various analyses and applications of logically constrained term rewriting. A fundamental question on constrained terms arising there is how to characterize equivalence between th

  17. Naoto Yoshida, Tadahiro Taniguchi

    In multi-agent reinforcement learning (MARL), effective communication improves agent performance, particularly under partial observability. We propose MARL-CPC, a framework that enables communication among fully decentralized, independent agents without parameter sharing. MARL-CPC incorporates a message learning model based on collective predictive coding (C

  18. Dongwei Wang, Zijie Liu, Song Wang, Yuxin Ren

    The Key-Value (KV) cache reading latency increases significantly with context lengths, hindering the efficiency of long-context LLM inference. To address this, previous works propose retaining a small fraction of KV cache based on token importance. For example, KV eviction uses static heuristics to retain tokens, while KV retrieval dynamically selects query-

  19. Ken Okamura, Yosuke Sato, Satoshi Takada

    This paper investigates the stress and displacement distribution in a two-dimensional elastic hollow disk subjected to distributed diametric loading, extending our previous analysis of concentrated loading [Okamura et al. Strength Mater. 57, 102-114 (2025)]. The study provides deeper insights into the mechanical behavior of materials such as concrete and roc

  20. M. M. López-Gutiérrez, H. Bravo-Alfaro, P. T. Rahna, G. A. Mamon

    During the fall of late-type galaxies into clusters, they can experiment a variety of evolutionary mechanisms according their local environment. Consequently, studying the UV emission and the cold gas of late-type galaxies provide key insights in the evolution of short-lived starburst and galaxy quenching. In this work, we conduted a study of two 28' fields

  21. Aaditya Shankar Majumder

    User experience research often uses surveys and interviews, which may miss subconscious user interactions. This study explores eye-tracking and biometric feedback as tools to assess user engagement and cognitive load in digital interfaces. These methods measure gaze behavior and bodily responses, providing an objective complement to qualitative insights. Usi

  22. Weiyu Liu, Neil Nie, Ruohan Zhang, Jiayuan Mao

    We introduce Behavior from Language and Demonstration (BLADE), a framework for long-horizon robotic manipulation by integrating imitation learning and model-based planning. BLADE leverages language-annotated demonstrations, extracts abstract action knowledge from large language models (LLMs), and constructs a library of structured, high-level action represen

  23. Hilay Shah, Freeke van de Voort, Amit Seta, Christoph Federrath

    We study gas mixing in a simulated Milky Way-mass galaxy's circumgalactic medium (CGM) using cosmological `zoom-in' simulations. We insert tracer dyes in the CGM with different gas flows (shearing, coherent, and static) and diverse physical properties to track gas mixing. We correlate the extent and shape of the dye spread with the local gas properties to un

  24. Fakhraddin Alwajih, Samar M. Magdy, Abdellah El Mekki, Omer Nacar

    Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. Constructed through advanced agentic workflows and extensive human-in-the-loop

  25. Wanfu Gao, Zengyao Man, Zebin He, Yuhao Tang

    Feature generation is a critical step in machine learning, aiming to enhance model performance by capturing complex relationships within the data and generating meaningful new features. Traditional feature generation methods heavily rely on domain expertise and manual intervention, making the process labor-intensive and challenging to adapt to different scen

  26. Khoa Ta

    Using the technique of inductive resolution introduced in arXiv:2303.07979, we prove that the homology of Rook-Brauer Algebra, interpreted as appropriate Tor-group, is isomorphic to that of symmetric group for all degrees under the assumption that $\epsilon$ in $R$ is invertible; furthermore, we also prove the homology of the Motzkin algebras vanishes in pos

  27. Fuxin Guan, Zemeng Lin, Sixin Chen, Xinhua Wen

    Metamaterials exhibit extraordinary properties yet suffer from pronounced wave dissipation, particularly in optical imaging and sensing systems. Recent advances leveraging complex frequency wave excitations with virtual gain effect, synthesized by multi-monochromatic waves, offer promising solutions for optical loss compensation. However, this approach faces

  28. Weiguang Zhang, Huangcheng Lu, Maizhen Ning, Xiaowei Huang

    Document dewarping aims to rectify deformations in photographic document images, thus improving text readability, which has attracted much attention and made great progress, but it is still challenging to preserve document structures. Given recent advances in diffusion models, it is natural for us to consider their potential applicability to document dewarpi

  29. Yu-Heng Hung, Kai-Jie Lin, Yu-Heng Lin, Chien-Yi Wang

    Bayesian optimization (BO) offers an efficient pipeline for optimizing black-box functions with the help of a Gaussian process prior and an acquisition function (AF). Recently, in the context of single-objective BO, learning-based AFs witnessed promising empirical results given its favorable non-myopic nature. Despite this, the direct extension of these appr

  30. Linyu Li, Zhi Jin, Yichi Zhang, Dongming Jin

    Knowledge graphs (KGs) play a key role in promoting various multimedia and AI applications. However, with the explosive growth of multi-modal information, traditional knowledge graph completion (KGC) models cannot be directly applied. This has attracted a large number of researchers to study multi-modal knowledge graph completion (MMKGC). Since MMKG extends

  31. Patrick Vossler, Fan Xia, Yifan Mai, Adarsh Subbaswamy

    Given the challenge of automatically evaluating free-form outputs from large language models (LLMs), an increasingly common solution is to use LLMs themselves as the judging mechanism, without any gold-standard scores. Implicitly, this practice accounts for only sampling variability (aleatoric uncertainty) and ignores uncertainty about judge quality (epistem

  32. Robert W. Heath,, Joseph Carlson, Nitish Vikas Deshpande, Miguel Rodrigo Castellanos

    We present an evolution of multiple-input multiple-output (MIMO) wireless communications known as the tri-hybrid MIMO architecture. In this framework, the traditional operations of linear precoding at the transmitter are distributed across digital beamforming, analog beamforming, and reconfigurable antennas. Compared with the hybrid MIMO architecture, which

  33. Wei Li, Mengcheng Lan, Jiaxing Xu, Yiping Ke

    Graphs are essential for modeling complex interactions across domains such as social networks, biology, and recommendation systems. Traditional Graph Neural Networks, particularly Message Passing Neural Networks (MPNNs), rely heavily on supervised learning, limiting their generalization and applicability in label-scarce scenarios. Recent self-supervised appr

  34. Bikai Gao

    We investigate the constraints on the strength of first-order phase transitions in neutron star matter and its relation to the origin of nucleon mass. By combining the parity doublet model for the hadronic phase, the Nambu-Jona-Lasinio model for quark matter, and the integral constraint framework for intermediate densities, we construct equation of states sp

  35. Tianjun Gu, Linfeng Li, Xuhong Wang, Chenghua Gong

    Adaptive navigation in unfamiliar environments is crucial for household service robots but remains challenging due to the need for both low-level path planning and high-level scene understanding. While recent vision-language model (VLM) based zero-shot approaches reduce dependence on prior maps and scene-specific training data, they face significant limitati

  36. Hyejeong Ryu

    Sampling-based motion planners such as Rapidly-exploring Random Tree* (RRT*) and its informed variant IRRT* are widely used for optimal path planning in complex environments. However, these methods often suffer from slow convergence and high variance due to their reliance on random sampling, particularly when initial solution discovery is delayed. This paper

  37. Juan Ren, Mark Dras, Usman Naseem

    Large Vision-Language Models (LVLMs) have shown remarkable capabilities across a wide range of multimodal tasks. However, their integration of visual inputs introduces expanded attack surfaces, thereby exposing them to novel security vulnerabilities. In this work, we conduct a systematic representational analysis to uncover why conventional adversarial attac

  38. Aditya Gunturu, Ben Pearman, Keiichi Ihara, Morteza Faraji

    We introduce MapStory, an LLM-powered animation prototyping tool that generates editable map animation sequences directly from natural language text by leveraging a dual-agent LLM architecture. Given a user written script, MapStory automatically produces a scene breakdown, which decomposes the text into key map animation primitives such as camera movements,

  39. Guo-Zhao Liao, Xiao-Feng Gong, Wei Liu, Hing Cheung So

    This paper investigates target localization using a multistatic multiple-input multiple-output (MIMO) radar system with two distinct coprime array configurations: coprime L-shaped arrays and coprime planar arrays. The observed signals are modeled as tensors that admit a coupled canonical polyadic decomposition (C-CPD) model. For each configuration, a C-CPD m

  40. Ziyun Zhang, Xinyi Liu, Xiaoyi Zhang, Jun Wang

    External knowledge has played a crucial role in the recent development of computer use agents. We identify a critical knowledge-execution gap: retrieved knowledge often fails to translate into effective real-world task execution. Our analysis shows even 90% correct knowledge yields only 41% execution success rate. To bridge this gap, we propose UI-Evol, a pl

  41. Taro Yano, Yoichi Ishibashi, Masafumi Oyamada

    Large Language Models (LLMs) have demonstrated exceptional performance across a wide range of tasks. To further tailor LLMs to specific domains or applications, post-training techniques such as Supervised Fine-Tuning (SFT), Preference Learning, and model merging are commonly employed. While each of these methods has been extensively studied in isolation, the

  42. Mengjingcheng Mo, Xinyang Tong, Mingpi Tan, Jiaxu Leng

    While unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex scenes. Existing datasets and methods, mainly designed for fixed ground-level views, struggle to adapt to these conditions, leading to significant performance drops in drone-view s

  43. S. V. Mousavi

    This study examines a system of three coupled qubits, focusing on entanglement measures in the presence of decoherence. It utilizes an XXZ Heisenberg chain with an external magnetic field and Dzyaloshinskii-Moriya interaction, considering intrinsic decoherence. The results reveal that only the magnetic field strength affects entanglement, while intrinsic dec

  44. Senmao Li, Lei Wang, Kai Wang, Tao Liu

    Text-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference speed and image quality, posing challenges for efficient deployment. Existing distilled T2I models can generate high-fidelity images with fewer sampling steps, but often struggle with diversity and quality, especiall

  45. Aakriti Agrawal, Mucong Ding, Zora Che, Chenghao Deng

    With Large Language Models (LLMs) rapidly approaching and potentially surpassing human-level performance, it has become imperative to develop approaches capable of effectively supervising and enhancing these powerful models using smaller, human-level models exposed to only human-level data. We address this critical weak-to-strong (W2S) generalization challen

  46. Qihuang Zhong, Liang Ding, Fei Liao, Juhua Liu

    Domain-specific instruction-tuning has become the defacto standard for improving the performance of large language models (LLMs) in specialized applications, e.g., medical question answering. Since the instruction-tuning dataset might contain redundant or low-quality data, data selection (DS) is usually required to maximize the data efficiency. Despite the s

  47. Jinxiao Qin, Hong-Liang Yan, Wenyuan Cui, Jian-Rong Shi

    Whether the presence of planets affects the lithium (Li) abundance of their host stars is still an open question. To investigate the difference of the Li abundance between planet-host stars (HS) and isolated stars (IS) with no detected planets, we analyze a large sample of stars with temperatures ranging from 4600 to 6600 K and metallicity ranging from -0.55

  48. Mengdan Zhu, Senhao Cheng, Guangji Bai, Yifei Zhang

    Text-to-image generation increasingly demands access to domain-specific, fine-grained, and rapidly evolving knowledge that pretrained models cannot fully capture, necessitating the integration of retrieval methods. Existing Retrieval-Augmented Generation (RAG) methods attempt to address this by retrieving globally relevant images, but they fail when no singl

  49. Insu Lee, Wooje Park, Jaeyun Jang, Minyoung Noh

    Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted cameras serves as key input. While this view offers fine-grained cues about user attention and hand-object interactions, its narrow field of view and lack of global c

  50. Kavita Kumari, I. E. Papadakis, G. C. Dewangan

    We conducted a comprehensive spectral and timing analysis of NGC 6814 using AstroSat's 2019 and XMM-Newton's 2021 observations. Cross-correlation analysis revealed a significant correlation between FUV (1541 \AA)/X-ray and UVW1 (2910 \AA)/X-ray variations, with delays of $\sim 15~\rm{ks}$ and $30~\rm{ks}$, respectively. We constructed four broadband SEDs aft

  51. Mai Ali, Christopher Lucasius, Tanmay P. Patel, Madison Aitken

    Speech is a noninvasive digital phenotype that can offer valuable insights into mental health conditions, but it is often treated as a single modality. In contrast, we propose the treatment of patient speech data as a trimodal multimedia data source for depression detection. This study explores the potential of large language model-based architectures for sp

  52. Masahiko Ueda

    Controlling payoffs in repeated games is one of the important topics in control theory of multi-agent systems. Recently proposed zero-determinant strategies enable players to unilaterally enforce linear relations between payoffs. Furthermore, based on the mathematics of zero-determinant strategies, regional payoff control, in which payoffs are enforced into

  53. Zijian Yang, Yulin Shao, Shaodan Ma

    Energy consumption and device lifetime are critical concerns for battery-constrained IoT devices. This paper introduces the Feedback-Aided Coding and Energy Transfer (FACET) framework, which synergistically combines adaptive feedback channel coding with wireless power transfer. FACET leverages the saturation effect of feedback coding, where increasing downli

  54. Long Tan Le, Senura Hansaja Wanasekara, Zerun Niu, Nguyen H. Tran

    Semantic communication (SemCom) has emerged as a promising paradigm for 6G wireless systems by transmitting task-relevant information rather than raw bits, yet existing approaches remain vulnerable to dual sources of uncertainty: semantic misinterpretation arising from imperfect feature extraction and transmission-level perturbations from channel noise. Curr

  55. Joshua Rooney

    For a positive integer $n$, an $n$-tuple of dice $(A_1,A_2,\dots,A_n)$ is called balanced if $P(A_1<A_2) = P(A_2<A_3) = \cdots = P(A_n<A_1)$ and nontransitive if $P(A_1<A_2), P(A_2<A_3), \dots, P(A_n<A_1)$ are each greater than $\frac{1}{2}$. For a balanced and nontransitive $n$-tuple of dice $(A_1,A_2,\dots,A_n)$, we define the winning probability $w(A_1,A_

  56. Jiayi Liu, Jonathan E. Ron, Giulia Rinaldi, Ivanna Williantarra

    Cell migration in vivo is often guided by chemical signals. Such chemotaxis, such as performed by immune cells migrating to a wound site, is complicated by the complex geometry inside living tissues. In this study, we extend our theoretical model of branched-cell migration on a network by introducing chemokine sources to explore the cellular response. The mo

  57. L. J. M. Davies, J. E. Thorne, S. Bellstedt, R. H. W. Cook

    In part I of this series we discussed the variation of star-formation histories (SFHs) across the specific star formation rate - stellar mass plane (sSFR-M$_{\star}$) using the Deep Extragalactic VIsible Legacy Survey (DEVILS). Here we explore the physical mechanisms that are likely driving these observational trends, by comparing the properties of galaxies

  58. L. J. M. Davies, J. E. Thorne, S. Bellstedt, R. H. W. Cook

    In a recent paper we parameterised the evolution of the star-formation rate dispersion ($\sigma_{SFR}$) across the specific star-formation rate - stellar mass plane (sSFR-M$_{\star}$) using the Deep Extragalactic VIsible Legacy Survey (DEVILS) - suggesting that the point at which the minimum in the dispersion occurs (M$^{*}_{\sigma-min}$) defines a boundary

  59. Sinan Wang, Junwei Zhou, Fan Feng, Zhiqi Li

    We propose the Vortex Particle Flow Map (VPFM) method to simulate incompressible flow with complex vortical evolution in the presence of dynamic solid boundaries. The core insight of our approach is that vorticity is an ideal quantity for evolution on particle flow maps, enabling significantly longer flow map distances compared to other fluid quantities like

  60. Maria. S. Kirsanova, Anastasiia A. Farafontova

    Since the emission of water molecules cannot be observed from Earth, less abundant isotopologues, such as H$_2^{18}$O and HDO, are used to trace water in star-forming regions. The main aim of this study is to determine HDO abundance in the hot core RCW 120 S2. We performed observations of the hot core in the 200-255~GHz range using the nFLASH230 receiver on

  61. Linli Zhou, Bokun Wang, My T. Thai, Tianbao Yang

    Two-way partial AUC (TPAUC) is a critical performance metric for binary classification with imbalanced data, as it focuses on specific ranges of the true positive rate (TPR) and false positive rate (FPR). However, stochastic algorithms for TPAUC optimization remain under-explored, with existing methods either limited to approximated TPAUC loss functions or b

  62. Wei Lin, Chenyang Zhao, Antoni B. Chan

    Point detection has been developed to locate pedestrians in crowded scenes by training a counter through a point-to-point (P2P) supervision scheme. Despite its excellent localization and counting performance, training a point-based counter still faces challenges concerning annotation labor: hundreds to thousands of points are required to annotate a single sa

  63. Prashant Bhat, Laurens Niesten, Elahe Arani, Bahram Zonooz

    Continual learning (CL) has remained a significant challenge for deep neural networks as learning new tasks erases previously acquired knowledge, either partially or completely. Existing solutions often rely on experience rehearsal or full model surrogates to mitigate CF. While effective, these approaches introduce substantial memory and computational overhe

  64. Elangbam Chingkheinganba Meetei, S. Surendra Singh

    In the present article, we developed a dynamical system in the context of modified $f(R,G,T)$ gravity, where $R$, $G$ and $T$ are Ricci scalar, Gauss-Bonnet term and energy-momentum tensor respectively. Development of the dynamical system is done by first defining 9 dimensionless variables and formulate a ordinary differential equations by taking derivative

  65. Ashim Gupta, Vivek Srikumar

    Inference-time scaling via repeated sampling has shown promise in reasoning tasks, but its effectiveness in multilingual generation remains underexplored. We evaluate this approach using perplexity- and reward-based verifiers on two multilingual benchmarks: the Aya Evaluation Suite and m-ArenaHard. Our results show consistent quality improvements, with gains

  66. Bolei He, Xinran He, Mengke Chen, Xianwei Xue

    Large Language Models (LLMs) excel in many areas but continue to face challenges with complex reasoning tasks, such as Multi-Hop Question Answering (MHQA). MHQA requires integrating evidence from diverse sources while managing intricate logical dependencies, often leads to errors in reasoning. Retrieval-Augmented Generation (RAG), widely employed in MHQA tas

  67. Chenglin Fan, Dahoon Lee, Euiwoong Lee

    Correlation Clustering (CC) is a foundational problem in unsupervised learning that models binary similarity relations using labeled graphs. While classical CC has been widely studied, many real-world applications involve more nuanced relationships, either multi-class categorical interactions or varying confidence levels in edge labels. To address these, two

  68. Qirun Zeng, Eric He, Richard Hoffmann, Xuchuang Wang

    Adversarial attacks on stochastic bandits have traditionally relied on some unrealistic assumptions, such as per-round reward manipulation and unbounded perturbations, limiting their relevance to real-world systems. We propose a more practical threat model, Fake Data Injection, which reflects realistic adversarial constraints: the attacker can inject only a

  69. Yoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi

    Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still struggle with fine-grained text-image alignment, often failing to faithfully depict objects with correct attributes such as color, shape, and spatial relations. To mitigate this issue, previous studies have explored pr

  70. Pratik Rakesh Singh, Kritarth Prasad, Mohammadi Zaki, Pankaj Wasnik

    Translating multi-word expressions (MWEs) and idioms requires a deep understanding of the cultural nuances of both the source and target languages. This challenge is further amplified by the one-to-many nature of idiomatic translations, where a single source idiom can have multiple target-language equivalents depending on cultural references and contextual v

  71. Zeyi Liao, Jaylen Jones, Linxi Jiang, Yuting Ning

    Computer-use agents (CUAs) promise to automate complex tasks across operating systems (OS) and the web, but remain vulnerable to indirect prompt injection. Current evaluations of this threat either lack support realistic but controlled environments or ignore hybrid web-OS attack scenarios involving both interfaces. To address this, we propose RedTeamCUA, an

  72. Kaiyu He, Zhiyu Chen

    Since the advent of Large Language Models (LLMs), efforts have largely focused on improving their instruction-following and deductive reasoning abilities, leaving open the question of whether these models can truly discover new knowledge. In pursuit of artificial general intelligence (AGI), there is a growing need for models that not only execute commands or

  73. Sarah A. Vollert, Christopher Drovandi, Cailan Jeynes-Smith, Luz V. Pascal

    Mathematical models connect theory with the real world through data, enabling us to interpret, understand, and predict complex phenomena. However, scientific knowledge often extends beyond what can be empirically measured, offering valuable insights into complex and uncertain systems. Here, we introduce a statistical framework for calibrating mathematical mo

  74. R. S. Watson, K. V. Kheruntsyan

    We apply a simple sudden quench approximation for the unitary work strokes of a quantum Otto engine in order to provide a general analysis of its performance, applicable to arbitrary quantum models with two-body interactions. This work extends recent results for an interaction-driven Otto cycle to generic many-body interacting quantum models, providing unive

  75. Adriana L. Duncan, Joe Kileel

    Group synchronization is the problem of determining reliable global estimates from noisy local measurements on networks. The typical task for group synchronization is to assign elements of a group to the nodes of a graph in a way that respects group elements given on the edges which encode information about local pairwise relationships between the nodes. In

  76. Sina Mohammadi, Ali Hassan, Rouzbeh Haghighi, Van-Hai Bui

    This paper investigates the capability of off-the-shelf large language models (LLMs) to solve the economic dispatch (ED) problem. ED is a hard-constrained optimization problem solved on a day-ahead timescale by grid operators to minimize electricity generation costs while accounting for physical and engineering constraints. Numerous approaches have been prop

  77. Tianxiang Zhan, Ming Jin, Yuanpeng He, Yuxuan Liang

    Recurring concept drift is pervasive in real-world online time series, where the underlying data-generating process repeatedly alternates between a small set of regimes, most notably daily or seasonal cycles that dominate energy, traffic, and weather patterns, and is therefore a central obstacle to reliable long-horizon forecasting. This problem poses a dual

  78. Dongyue Li, Ziniu Zhang, Lu Wang, Hongyang R. Zhang

    This paper develops an ensemble method for fine-tuning a language model to multiple datasets. Existing methods, such as quantized LoRA (QLoRA), are efficient when adapting to a single dataset. When training on multiple datasets of different tasks, a common setup in practice, it remains unclear how to design an efficient adaptation for fine-tuning language mo

  79. Linhui Wu, Fu-Guo Xie, Qian Zheng, Quan Guo

    This study investigates the projected, quasi-symmetric $\sim\rm46\,kpc$-scale diffuse radio lobes surrounding the giant elliptical galaxy M\,87, utilizing well-sampled wideband ($\rm 60\,MHz-10.55\,GHz$) observations from MWA and VLA, supplemented by data from LOFAR and Effelsberg. The observed structures feature sharp edges and filaments, with nearly unifor

  80. Lianghui Zhu, Xitong Ling, Minxi Ouyang, Xiaoping Liu

    Gastrointestinal (GI) diseases represent a clinically significant burden, necessitating precise diagnostic approaches to optimize patient outcomes. Conventional histopathological diagnosis suffers from limited reproducibility and diagnostic variability. To overcome these limitations, we develop Digepath, a specialized foundation model for GI pathology. Our f

  81. Hong-Lin Lin, Pragati Aashna, Yu Cao, Aaron Danner

    Electro-optic modulators are indispensable components of modern day photonic integrated circuits (PICs). Recently lithium niobate has emerged as a key material to realize large-bandwidth high-speed modulation, but next-generation modulators require high-density integration, low cost, low power and high performance simultaneously, which are difficult to achie

  82. Yin Hua, Zhiqiang Liu, Mingyang Chen, Zheng Fang

    In natural language processing (NLP) and computer vision (CV), the successful application of foundation models across diverse tasks has demonstrated their remarkable potential. However, despite the rich structural and textual information embedded in knowledge graphs (KGs), existing research of foundation model for KG has primarily focused on their structural

  83. Chong Zeng, Yue Dong, Pieter Peers, Hongzhi Wu

    We present RenderFormer, a neural rendering pipeline that directly renders an image from a triangle-based representation of a scene with full global illumination effects and that does not require per-scene training or fine-tuning. Instead of taking a physics-centric approach to rendering, we formulate rendering as a sequence-to-sequence transformation where

  84. Kamyar Rashidi, Matthew Y. Sfeir

    Terahertz Time-Domain Spectroscopy is a powerful technique for extracting the low-frequency optical properties of materials. However, the optical constants are difficult to determine directly from the experimental transfer function, such that various numerical approximations must be implemented to describe specific conditions. Here, we introduce a modified T

  85. Asal Mehradfar, Xuzhe Zhao, Yilun Huang, Emir Ceyani

    Designing analog circuits from performance specifications is a complex, multi-stage process encompassing topology selection, parameter inference, and layout feasibility. We introduce FALCON, a unified machine learning framework that enables fully automated, specification-driven analog circuit synthesis through topology selection and layout-constrained optimi

  86. Juseung Oh, Wontaek Kim, Gyouil Jeong, Yeri Lee

    Optical second-harmonic generation (SHG) enables orientational polarimetry for crystallographic analysis and domain imaging of various materials. However, conventional intensity polarimetry, which neglects phase information, fails to resolve antiparallel domains and to describe two-dimensional heterostructures, which represent a new class of van der Waals-bo

  87. Yuanhong Zhang, Muyao Yuan, Weizhan Zhang, Tieliang Gong

    The Segment Anything Model (SAM), a vision foundation model, exhibits impressive zero-shot capabilities in general tasks but struggles in specialized domains. Parameter-efficient fine-tuning (PEFT) is a promising approach to unleash the potential of SAM in novel scenarios. However, existing PEFT methods for SAM neglect the domain-invariant relations encoded

  88. Yue Zhu, Hao Yu, Chen Wang, Zhuoran Liu

    The increasing adoption of large language models (LLMs) with extended context windows necessitates efficient Key-Value Cache (KVC) management to optimize inference performance. Inference workloads like Retrieval-Augmented Generation (RAG) and agents exhibit high cache reusability, making efficient caching critical to reducing redundancy and improving speed.

  89. Haruki Kai, Tsuyoshi Okita

    We developed a deep learning algorithm for human activity recognition using sensor signals as input. In this study, we built a pretrained language model based on the Transformer architecture, which is widely used in natural language processing. By leveraging this pretrained model, we aimed to improve performance on the downstream task of human activity recog

  90. James Demmel, Ioana Dumitriu, Ryan Schneider

    This paper presents a fast, randomized divide-and-conquer algorithm for the definite generalized eigenvalue problem, which corresponds to pencils $(A,B)$ in which $A$ and $B$ are Hermitian and the Crawford number $γ(A,B) = \min_{\|x\|_2 = 1} |x^H(A+iB)x|$ is positive. Adapted from the fastest known method for diagonalizing arbitrary matrix pencils [Foundatio

  91. Jialong Guo, Xinghao Chen, Yehui Tang, Yunhe Wang

    Large language models(LLMs) have garnered significant attention and demonstrated impressive capabilities in a wide range of applications. However, due to their enormous computational costs, the deployment and application of LLMs are often severely limited. To address this issue, structured pruning is an effective solution to compress the parameters of LLMs.

  92. Mir Sazzat Hossain, Ovi Paul, Md Akil Raihan Iftee, Rakibul Hasan Rajib

    Land Use Land Cover (LULC) mapping using deep learning significantly enhances the reliability of LULC classification, aiding in understanding geography, socioeconomic conditions, poverty levels, and urban sprawl. However, the scarcity of annotated satellite data, especially in South/East Asian developing countries, poses a major challenge due to limited fund

  93. Chenfeng Wei, Qi Wu, Si Zuo, Jiahua Xu

    Autonomous driving datasets are essential for validating the progress of intelligent vehicle algorithms, which include localization, perception, and prediction. However, existing datasets are predominantly focused on structured urban environments, which limits the exploration of unstructured and specialized scenarios, particularly those characterized by sign

  94. Mengqi Zhang

    Time-dependent fluid dynamics plays a crucial role in both natural phenomena and industrial applications. Understanding the flow instabilities and transitions within these dynamical systems is essential for predicting and controlling their unsteady behaviour. A classic example of time-dependent flow is the Stokes layer. To study the transition mechanism in t

  95. Marvin Limpijankit, John Kender

    We propose a two-step approach for detecting differences in the style of images across sources of differing cultural affinity, where images are first clustered into finer visual themes based on content before their aesthetic features are compared. We test this approach on 2,400 YouTube video thumbnails taken equally from two U.S. and two Chinese YouTube chan

  96. Yiheng Lin, Shifang Zhao, Ting Liu, Xiaochao Qu

    Personalized image generation aims to integrate user-provided concepts into text-to-image models, enabling the generation of customized content based on a given prompt. Recent zero-shot approaches, particularly those leveraging diffusion transformers, incorporate reference image information through multi-modal attention mechanism. This integration allows the

  97. Xianbiao Qi, Yelin He, Jiaquan Ye, Chun-Guang Li

    Scaling Transformer to a large scale without using some technical tricks such as learning rate warump and using an obviously lower learning rate is an extremely challenging task, and is increasingly gaining more attention. In this paper, we provide a theoretical analysis for the process of training Transformer and reveal the rationale behind the model crash

  98. Jiawei Fu, Donald P. Green

    How should researchers analyze randomized experiments in which the main outcome is latent and measured in multiple ways but each measure contains some degree of error? We first identify a critical study-specific noncomparability problem in existing methods for handling multiple measurements, which often rely on strong modeling assumptions or arbitrary standa

  99. Hanyin Wang, Zhenbang Wu, Gururaj Kolar, Hariprasad Korsapati

    Diagnosis-Related Group (DRG) codes are essential for hospital reimbursement and operations but require labor-intensive assignment. Large Language Models (LLMs) struggle with DRG coding due to the out-of-distribution (OOD) nature of the task: pretraining corpora rarely contain private clinical or billing data. We introduce DRG-Sapphire, which uses large-scal

  100. Saleh Afzoon, Ali Shahsavandi, Phuong Thao Huynh, Melika Zare

    AI copilots represent a new generation of AI-powered systems designed to assist users, particularly knowledge workers and developers, in complex, context-rich tasks. As these systems become more embedded in daily workflows, personalization has emerged as a critical factor for improving usability, effectiveness, and user satisfaction. Central to this personal