Skip to content

May 2025 arXiv papers — page 91

Showing 9,0019,100 of 24,552 papers

  1. Vijeta Deshpande, Debasmita Ghose, John D. Patterson, Roger Beaty

    Diverse language model responses are crucial for creative generation, open-ended tasks, and self-improvement training. We show that common diversity metrics, and even reward models used for preference optimization, systematically bias models toward shorter outputs, limiting expressiveness. To address this, we introduce Diverse, not Short (Diverse-NS), a leng

  2. Masanari Kimura, Howard Bondell

    The power prior is a class of informative priors designed to incorporate historical data alongside current data in a Bayesian framework. It includes a power parameter that controls the influence of historical data, providing flexibility and adaptability. A key property of the power prior is that the resulting posterior minimizes a linear combination of KL di

  3. James A. Rossmanith, Christine Vaughan

    Vlasov equations model the dynamics of plasma in the collisionless regime. A standard approach for numerically solving the Vlasov equation is to operator split the spatial and velocity derivative terms, allowing simpler time-stepping schemes to be applied to each piece separately (known as the Cheng-Knorr method). One disadvantage of such an operator split m

  4. Runze Yan, Xun Shen, Akifumi Wachi, Sebastien Gros

    When applying offline reinforcement learning (RL) in healthcare scenarios, the out-of-distribution (OOD) issues pose significant risks, as inappropriate generalization beyond clinical expertise can result in potentially harmful recommendations. While existing methods like conservative Q-learning (CQL) attempt to address the OOD issue, their effectiveness is

  5. Yuhan Ji, Song Gao, Ying Nie, Ivan Majić

    Applying AI foundation models directly to geospatial datasets remains challenging due to their limited ability to represent and reason with geographical entities, specifically vector-based geometries and natural language descriptions of complex spatial relations. To address these issues, we investigate the extent to which a well-known-text (WKT) representati

  6. Viet-Anh Nguyen, Shiqian Zhao, Gia Dao, Runyi Hu

    Recently, Large Reasoning Models (LRMs) have demonstrated superior logical capabilities compared to traditional Large Language Models (LLMs), gaining significant attention. Despite their impressive performance, the potential for stronger reasoning abilities to introduce more severe security vulnerabilities remains largely underexplored. Existing jailbreak me

  7. Camelia R. Walker, Md Nurul Anwar, Leandra Brauninger, Jack Richards

    Plasmodium falciparum is responsible for the majority of malaria morbidity and mortality each year. Malaria transmission rates vary by location and time of year due to climate and environmental conditions. We show the impact of these factors by developing a stochastic spatiotemporal agent-based malaria model that captures the impact of spatially distributed

  8. Zheng Chen, Zichen Zou, Kewei Zhang, Xiongfei Su

    Diffusion models have demonstrated promising performance in real-world video super-resolution (VSR). However, the dozens of sampling steps they require, make inference extremely slow. Sampling acceleration techniques, particularly single-step, provide a potential solution. Nonetheless, achieving one step in VSR remains challenging, due to the high training o

  9. Chandan Kumar, Amrit Kumar Mondal, Sreya Pal, Sayan Mathur

    Three-dimensional (3D) magnetic nanostructures offer a versatile platform for exploring complex spin textures and spin-wave (SW) dynamics, with implications in next-generation spintronic and magnonic technologies. Advances in 3D nanofabrication have allowed a wide-range of structures and phenomena to be realized. Whilst the study of simple cylindrical magnet

  10. Derong Xu, Pengyue Jia, Xiaopeng Li, Yingyi Zhang

    Despite the strong abilities, large language models (LLMs) still suffer from hallucinations and reliance on outdated knowledge, raising concerns in knowledge-intensive tasks. Graph-based retrieval-augmented generation (GRAG) enriches LLMs with knowledge by retrieving graphs leveraging relational evidence, but it faces two challenges: structure-coupled irrele

  11. Kaiyue Hou, Shuowen Zhang

    This paper studies a networked sensing system with multiple base stations (BSs), which collaboratively sense the unknown and random three-dimensional (3D) location of a target based on the target-reflected echo signals received at the BSs. Considering a practical scenario where the target location distribution is known a priori for exploitation, we aim to de

  12. Rashed Shelim, Shengzhe Xu, Walid Saad, Naren Ramakrishnan

    Vector representations of contextual embeddings learned by pre-trained large language models (LLMs) are effective in various downstream tasks in numerical domains such as time series forecasting. Despite their significant benefits, the tendency of LLMs to hallucinate in such domains can have severe consequences in applications such as energy, nature, finance

  13. Kosei Nakagawa, Yoshiko Kanada-En'yo

    The $0_2^+$ state of ${}^8\mathrm{He}$ has been discovered by a recent experiment, which suggested a developed cluster structure of spatially correlated neutron pairs, called "dineutrons" (${}^2n$). We aim to investigate the structure of ${}^8\mathrm{He}(0^+_1)$ and ${}^8\mathrm{He}(0^+_2)$ to clarify the monopole excitation mode in ${}^8\mathrm{He}$ system

  14. Wei Zhang, Zhenhong Zhou, Kun Wang, Junfeng Fang

    While large language models (LLMs) can solve PhD-level reasoning problems over long context inputs, they still struggle with a seemingly simpler task: following explicit length instructions-e.g., write a 10,000-word novel. Additionally, models often generate far too short outputs, terminate prematurely, or even refuse the request. Existing benchmarks focus p

  15. Rajesh Kumar, Suchi Kumari, Anubhav Mishra

    Real-world complex systems exhibit intricate interconnections and dependencies, especially social networks, technological infrastructures, and communication networks. These networks are prone to disconnection due to random failures or external attacks on their components. Therefore, managing the security and resilience of such networks is a prime concern, pa

  16. Ali Sarosh Bangash, Krish Veera, Ishfat Abrar Islam, Raiyan Abdul Baten

    An objective, face-valid method for scoring idea originality is to measure each idea's statistical infrequency within a population -- an approach long used in creativity research. Yet, computing these frequencies requires manually bucketing idea rephrasings, a process that is subjective, labor-intensive, error-prone, and brittle at scale. We introduce MuseSc

  17. Amber Young, Tyler Robinson, Joshua Krissansen-Totton, Edward Schwieterman

    Robust exoplanet characterization studies are underway, and the community is looking ahead toward developing observational strategies to search for life beyond our solar system. With the development of life detection approaches like searching for atmospheric chemical species indicative of life, chemical disequilibrium has also been proposed as a potentially

  18. Shuo Zheng, Shuowen Zhang

    Beyond diagonal intelligent reflecting surface (BD-IRS) is a new promising IRS architecture for which the reflection matrix is not limited to the diagonal structure as for conventional IRS. In this paper, we study a BD-IRS aided uplink integrated sensing and communication (ISAC) system where sensing is performed in a device-based manner. Specifically, we aim

  19. Yuren Mao, Wenyi Xu, Yuyang Qin, Yunjun Gao

    Computed Tomography (CT) scan, which produces 3D volumetric medical data that can be viewed as hundreds of cross-sectional images (a.k.a. slices), provides detailed anatomical information for diagnosis. For radiologists, creating CT radiology reports is time-consuming and error-prone. A visual question answering (VQA) system that can answer radiologists' que

  20. Wei-Lun Huang, Joshua Liu, Davood Tashayyod, Jun Kang

    Total Body Photography (TBP) is becoming a useful screening tool for patients at high risk for skin cancer. While much progress has been made, existing TBP systems can be further improved for automatic detection and analysis of suspicious skin lesions, which is in part related to the resolution and sharpness of acquired images. This paper proposes a novel sh

  21. Bohao Wu, Qingyun Wang, Yue Guo

    Personalizing jargon detection and explanation is essential for making technical documents accessible to readers with diverse disciplinary backgrounds. However, tailoring models to individual users typically requires substantial annotation efforts and computational resources due to user-specific finetuning. To address this, we present a systematic study of p

  22. Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou, Yu-Feng Li

    Tabular data, owing to its ubiquitous presence in real-world domains, has garnered significant attention in machine learning research. While tree-based models have long dominated tabular machine learning tasks, the recently proposed deep learning model TabPFN v2 has emerged, demonstrating unparalleled performance and scalability potential. Although extensive

  23. Zihan Chen, Song Wang, Zhen Tan, Jundong Li

    In-Context Learning (ICL) empowers Large Language Models (LLMs) to tackle diverse tasks by incorporating multiple input-output examples, known as demonstrations, into the input of LLMs. More recently, advancements in the expanded context windows of LLMs have led to many-shot ICL, which uses hundreds of demonstrations and outperforms few-shot ICL, which relie

  24. Chuan-Ren Chen, Cheng-Wei Chiang, Leon M. G. de la Vega

    In this work, we study a dark matter scenario where a dark Z boson possessing mass mixing with the SM Z boson couples to the DM candidate and serves as the portal to the SM. The UV origin of the mass mixing in the form of an extra dark Higgs doublet and a scalar dark singlet provides new exotic scalars which can constitute the final state of DM annihilation

  25. Sangyong Lee, Subo Hwang, Dohoon Kim

    In this paper, we propose MADCluster, a novel model-agnostic anomaly detection framework utilizing self-supervised clustering. MADCluster is applicable to various deep learning architectures and addresses the 'hypersphere collapse' problem inherent in existing deep learning-based anomaly detection methods. The core idea is to cluster normal pattern data into

  26. Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang

    With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations. While this offers scalability and flexibility, it also raises a critical, unresolved question: Can LLM judges fairly and robustly evaluate

  27. Yifan Zhang, Xinkui Zhao, Zuxin Wang, Guanjie Cheng

    The rapid advancement of large language models has unlocked remarkable capabilities across a diverse array of natural language processing tasks. However, the considerable differences among available LLMs-in terms of cost, performance, and computational demands-pose significant challenges for users aiming to identify the most suitable model for specific tasks

  28. Liang-Yeh Shen, Shi-Xin Fang, Yi-Cheng Lin, Huang-Cheng Chou

    This paper introduces Meta-PerSER, a novel meta-learning framework that personalizes Speech Emotion Recognition (SER) by adapting to each listener's unique way of interpreting emotion. Conventional SER systems rely on aggregated annotations, which often overlook individual subtleties and lead to inconsistent predictions. In contrast, Meta-PerSER leverages a

  29. Kotaro Miyake, Yasuhiro Yamaguchi

    The nature of the $X(3872)$ and other exotic hadrons has been a subject of extensive investigation since the first observation of the $X(3872)$ in 2003. While various theoretical models have been proposed, including hadronic molecular and compact tetraquark interpretations, some experimental evidence suggests that the $X(3872)$ may be a mixture state of a ha

  30. Yiliang Li, Ping Zhang, Zhengyuan Tian, Li Feng

    The H I Lyman-alpha (Ly$\alpha$) emission, with a wavelength of 1216 \r{A}, is the brightest solar ultraviolet (UV) line. However, comprehensive observations of the Ly$\alpha$ emission line across the full solar disk remain limited. As part of the ASO-S mission, the Solar Disk Imager (SDI) has successfully captured full-disk images in the Ly$\alpha$ band. Ga

  31. Hon Tik Tse, Siddarth Chandrasekar, Marlos C. Machado

    In recent years, the successor representation (SR) has attracted increasing attention in reinforcement learning (RL), and it has been used to address some of its key challenges, such as exploration, credit assignment, and generalization. The SR can be seen as representing the underlying credit assignment structure of the environment by implicitly encoding it

  32. Jisu Kim, Youngwoo Shin, Uiji Hwang, Jihun Choi

    Idioms have long posed a challenge due to their unique linguistic properties, which set them apart from other common expressions. While recent studies have leveraged large language models (LLMs) to handle idioms across various tasks, e.g., idiom-containing sentence generation and idiomatic machine translation, little is known about the underlying mechanisms

  33. Md Ashraf Uddin, Nam H. Chu, Reza Rafeh, Mutaz Barika

    Due to its nature of dynamic, mobility, and wireless data transfer, the Internet of Vehicles (IoV) is prone to various cyber threats, ranging from spoofing and Distributed Denial of Services (DDoS) attacks to malware. To safeguard the IoV ecosystem from intrusions, malicious activities, policy violations, intrusion detection systems (IDS) play a critical rol

  34. Henry X. Liu, Xintao Yan, Haowei Sun, Tinghan Wang

    Autonomous vehicles (AVs) have significantly advanced in real-world deployment in recent years, yet safety continues to be a critical barrier to widespread adoption. Traditional functional safety approaches, which primarily verify the reliability, robustness, and adequacy of AV hardware and software systems from a vehicle-centric perspective, do not sufficie

  35. Kazuyuki Yagasaki

    We study the Kuramoto model (KM) having random natural frequencies and defined on uniform graphs that may be complete, random dense or random sparse. The natural frequencies are assumed to be independent and identically distributed on a bounded interval. In the previous work, the corresponding continuum limit (CL) was proven to approximate stable motions in

  36. Anfeng Xu, Tiantian Feng, So Hyun Kim, Somer Bishop

    Automatic Speech Recognition (ASR) has recently shown remarkable progress, but accurately transcribing children's speech remains a significant challenge. Recent developments in Large Language Models (LLMs) have shown promise in improving ASR transcriptions. However, their applications in child speech including conversational scenarios are underexplored. In t

  37. Kai Li, Can Shen, Yile Liu, Jirui Han

    The rapid development and widespread adoption of Audio Large Language Models (ALLMs) demand rigorous evaluation of their trustworthiness. However, existing evaluation frameworks are primarily designed for text and fail to capture vulnerabilities introduced by the acoustic properties of audio. We find that significant trustworthiness risks in ALLMs arise from

  38. Zhihang Cai, Xingjun Zhang, Zhendong Tan, Zheng Wei

    Large Language Models (LLMs) have demonstrated remarkable proficiency across a wide range of tasks. However, LLMs often require larger batch sizes to enhance throughput or longer context lengths to meet task demands, which significantly increases the memory resource consumption of the Key-Value (KV) cache during inference, becoming a major bottleneck in LLM

  39. Shuchang Ye, Usman Naseem, Mingyuan Meng, Dagan Feng

    Medical Visual Question Answering (MedVQA) is crucial for enhancing the efficiency of clinical diagnosis by providing accurate and timely responses to clinicians' inquiries regarding medical images. Existing MedVQA models suffered from modality preference bias, where predictions are heavily dominated by one modality while overlooking the other (in MedVQA, us

  40. Anton Erofeev, Balasubramanya T. Nadiga, Ilya Timofeyev

    We apply Echo-State Networks to predict time series and statistical properties of the competitive Lotka-Volterra model in the chaotic regime. In particular, we demonstrate that Echo-State Networks successfully learn the chaotic attractor of the competitive Lotka-Volterra model and reproduce histograms of dependent variables, including tails and rare events.

  41. Kentaro Onda, Yosuke Kashiwagi, Emiru Tsunoo, Hayato Futami

    Recent studies have highlighted the potential of discrete tokens derived from self-supervised learning (SSL) models for various speech-related tasks. These tokens serve not only as substitutes for text in language modeling but also as intermediate representations for tasks such as automatic speech recognition (ASR). However, discrete tokens are typically obt

  42. Lucas O. Lima, Julián Faúndez, Natanael C. Costa, Raimundo R. dos Santos

    We investigate some thermodynamic and magnetic properties of the Hubbard model on two three-dimensional extensions of the Lieb lattice: the perovskite Lieb lattice (PLL) and the layered Lieb lattice (LLL). Using determinant quantum Monte Carlo (DQMC) simulations alongside Hartree-Fock and cluster mean-field theory (CMFT) approaches, we analyze how flat-band

  43. Naeem Budhwani, Mohammad Faghani, Hayden Richard

    Static Application Security Testing (SAST) enables organizations to detect vulnerabilities in code early; however, major SAST platforms do not include visual aids and present little insight on correlations between tainted data chains. We propose VIVID - Vulnerability Information Via Data flow - a novel method to extract and consume SAST insights, which is to

  44. Ichiro Hashimoto

    In this paper, we provide sufficient conditions of benign overfitting of fixed width leaky ReLU two-layer neural network classifiers trained on mixture data via gradient descent. Our results are derived by establishing directional convergence of the network parameters and classification error bound of the convergent direction. Our classification error bound

  45. Jesus Sanchez

    We provide a recipe for building explicit representations of the real Clifford algebras once an explicit family is given in dimensions $1$ through $4$. We further give an explicit construction of spin coordinate systems for a given real spinor module and use it to explicitly compute the parallel transport of spinor fields. We further highlight some novelties

  46. Indronil Bhattacharjee, Christabel Wayllace

    KnowledgeTracing (KT) involves predicting students' knowledge states based on their interactions with Intelligent Tutoring Systems (ITS). A key challenge is the cold start problem, accurately predicting knowledge for new students with minimal interaction data. Unlike prior work, which typically trains KT models on initial interactions of all students and tes

  47. Don N. Page

    Alvarez-Dominguez, Garay, Martin-Martinez, and Polo-Gomez have suggested that ``it is not possible to concentrate enough light to precipitate the formation of an event horizon. We argue that the dissipative quantum effects coming from the self-interaction of light (such as vacuum polarization) are enough to prevent any meaningful buildup of energy that could

  48. Ivan V. Morozov

    The basis of this work is a simple, extended corollary of Wilson's theorem. This corollary generates many more quotients than those already generated by Wilson's theorem, and it was of interest to derive how they relate to each other and build on the established properties of the original quotients. The most important results that were found were expressions

  49. Kavya Gaddipati

    This technical article explores comprehensive strategies for integrating Electrostatic Discharge (ESD) protection diodes and termination resistors in LowVoltage Differential Signaling (LVDS) designs. The article examines critical aspects of protection mechanisms, design considerations, impedance matching, and placement optimization techniques. Through detail

  50. Chaochen Gao, Xing Wu, Zijia Lin, Debing Zhang

    High-quality long-context instruction data is essential for aligning long-context large language models (LLMs). Despite the public release of models like Qwen and Llama, their long-context instruction data remains proprietary. Human annotation is costly and challenging, while template-based synthesis methods limit scale, diversity, and quality. We introduce

  51. Rikuhei Umemoto, Keisuke Fujii

    In many real-world complex systems, the behavior can be observed as a collection of discrete events generated by multiple interacting agents. Analyzing the dynamics of these multi-agent systems, especially team sports, often relies on understanding the movement and interactions of individual agents. However, while providing valuable snapshots, event-based po

  52. Takashi Nakagawa, Shin-ichi Takahero, Youhei Sasaki

    The possibility of the emergence of a stratified region in the uppermost part of the Earth's outer core with long-term magnetic field generation is assessed, taking into account uncertainties in the thermal conductivity of the Earth's core and the present-day heat flow across the core-mantle boundary (CMB). The radial structures of the Earth's outer core are

  53. Jesus Sanchez, Andres Franco Valiente

    Given a contact sub-Riemannian manifold one obtains a non-integrable splitting of the tangent bundle into the directions along the contact distribution and the Reeb field. We generalize the construction of the Bismut superconnection to this non-integrable setting and show that although singularities appear within the superconnection, if one extracts the fini

  54. Xuewu Lin, Tianwei Lin, Lichao Huang, Hongyu Xie

    A key challenge in robot manipulation lies in developing policy models with strong spatial understanding, the ability to reason about 3D geometry, object relations, and robot embodiment. Existing methods often fall short: 3D point cloud models lack semantic abstraction, while 2D image encoders struggle with spatial reasoning. To address this, we propose SEM

  55. Zhi Zhong, Akira Takahashi, Shuyang Cui, Keisuke Toyama

    Foley synthesis aims to synthesize high-quality audio that is both semantically and temporally aligned with video frames. Given its broad application in creative industries, the task has gained increasing attention in the research community. To avoid the non-trivial task of training audio generative models from scratch, adapting pretrained audio generative m

  56. M. O. Ajeesh, A. O. Scheie, Yu Liu, L. Keller

    Ce$_{2}$$M$Al$_{7}$Ge$_{4}$ ($M=$ Co, Ir, Ni or Pd) are heavy-fermion materials and host a variety of ground states ranging from magnetism to non-Fermi liquid behavior. The Co, Ir, and Ni members of the series undergo magnetic ordering with decreasing transition temperatures. In contrast, the Pd compound does not magnetically order down to 0.4 K and shows no

  57. Daiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma

    The advancements in Multimodal Large Language Models (MLLMs) have enabled various multimodal tasks to be addressed under a zero-shot paradigm. This paradigm sidesteps the cost of model fine-tuning, emerging as a dominant trend in practical application. Nevertheless, Multimodal Sentiment Analysis (MSA), a pivotal challenge in the quest for general artificial

  58. Chaoya Jiang, Yongrui Heng, Wei Ye, Han Yang

    Recently, reasoning-based MLLMs have achieved a degree of success in generating long-form textual reasoning chains. However, they still struggle with complex tasks that necessitate dynamic and iterative focusing on and revisiting of visual regions to achieve precise grounding of textual reasoning in visual evidence. We introduce \textbf{VLM-R$^3$} (\textbf{V

  59. Kentaro Onda, Keisuke Imoto, Satoru Fukayama, Daisuke Saito

    Recently, a method for synthesizing foreign-accented speech only with native speech data using discrete tokens obtained from self-supervised learning (SSL) models was proposed. Considering limited availability of accented speech data, this method is expected to make it much easier to simulate foreign accents. By using the synthesized accented speech as liste

  60. Navid Seidi, Satyaki Roy, Sajal Das

    Federated Learning (FL) holds great promise for digital health by enabling collaborative model training without compromising patient data privacy. However, heterogeneity across institutions, lack of sustained reputation, and unreliable contributions remain major challenges. In this paper, we propose a robust, peer-driven reputation mechanism for federated he

  61. Sophie Wu, Jan Philip Wahle, Saif M. Mohammad

    This paper is the first investigation of the connection between emotion, embodiment, and everyday language in a large sample of natural language data. We created corpora of body part mentions (BPMs) in online English text (blog posts and tweets). This includes a subset featuring human annotations for the emotions of the person whose body part is mentioned in

  62. Zirui He, Mingyu Jin, Bo Shen, Ali Payani

    Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but controlling their behavior reliably remains challenging, especially in open-ended generation settings. This paper introduces a novel supervised steering approach that operates in sparse, interpretable representation spaces. We employ s

  63. Guanghe Li, Junming Zhao, Shengjie Wang, Yang Gao

    Robotic insertion is a highly challenging task that requires exceptional precision in cluttered environments. Existing methods often have poor generalization capabilities. They typically function in restricted and structured environments, and frequently fail when the plug and socket are far apart, when the scene is densely cluttered, or when handling novel o

  64. Kaiwen Zhou, Xuandong Zhao, Gaowen Liu, Jayanth Srinivasa

    Large Reasoning Models (LRMs) introduce a new generation paradigm of explicitly reasoning before answering, leading to remarkable improvements in complex tasks. However, they pose great safety risks against harmful queries and adversarial attacks. While recent mainstream safety efforts on LRMs, supervised fine-tuning (SFT), improve safety performance, we fin

  65. Gregoire Fournier, György Turán

    Ehrenfeucht-Fra\"iss\'e (EF) games are a basic tool in finite model theory for proving definability lower bounds, with many applications in complexity theory and related areas. They have been applied to study various logics, giving insights on quantifier rank and other logical complexity measures. In this paper, we present an EF game to capture formula size

  66. K. Y. Liang, R . Z. Zhang, Z. F. Lin, Z. J. Li

    Nematicity and magnetism are prevalent orders in high transition temperature (Tc) superconductors, coexisting in the parent compound of most material families. Quantum fluctuations of nematicity or spin orders are both plausible candidates for mediating unconventional Cooper pairing. Identifying the sole effect of a nematic quantum critical point (QCP) on th

  67. Haiyun Huang, Xiu Zhang, Junzhi Zhu, Jianfa Zhao

    We employ time-resolved optical reflectivity to investigate the ultrafast dynamics of the charge-transfer gap (CTG) in a parent cuprate compound Ca$_2$CuO$_2$Cl$_2$ (CCOC). We observe a persistent photoinduced red shift of the CTG that lasts up to 1000 ps. The red shift during the slow decay after 10 ps can be well modeled by the localized picture, whereas i

  68. Kentaro Onda, Keisuke Imoto, Satoru Fukayama, Daisuke Saito

    In this study, we gained insight that contributes to achieving accent-robust ASR using only native speech data. In human perception of non-native speech, the phenomenon known as "interlanguage speech intelligibility benefit" (ISIB) is observed, where non-native listeners who share the native language with the speaker understand the speech better compared eve

  69. Mohammad Reza Taesiri, Brandon Collins, Logan Bolton, Viet Dac Lai

    Generative AI (GenAI) holds significant promise for automating everyday image editing tasks, especially following the recent release of GPT-4o on March 25, 2025. However, what subjects do people most often want edited? What kinds of editing actions do they want to perform (e.g., removing or stylizing the subject)? Do people prefer precise edits with predicta

  70. Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi

    Evaluating image captions requires cohesive assessment of both visual semantics and language pragmatics, which is often not entirely captured by most metrics. We introduce Redemption Score(RS), a novel hybrid framework that ranks image captions by triangulating three complementary signals: (1) Mutual Information Divergence (MID) for global image-text distrib

  71. Ilya I. Bogdanov, Elizaveta Neustroeva, Georgy Sokolov, Alexei Volostnov

    The paper is devoted to sufficient conditions for the existence of vertex cuts in simple graphs, where the induced subgraph on the cut vertices belongs to a specified graph class. In particular, we show that any connected graph with $n$ vertices and fewer than $(19n - 28)/8$ edges admits a forest cut. This result improves upon recent bounds, although it does

  72. Shuai Wang, Yizhou Sun, Judea Pearl, Ang Li

    Probabilities of causation play a central role in modern decision making. Tian and Pearl first introduced formal definitions and derived tight bounds for three binary probabilities of causation, such as the probability of necessity and sufficiency (PNS). However, estimating these probabilities requires both experimental and observational distributions specif

  73. Linfeng Qi, Zhaoyang Jia, Jiahao Li, Bin Li

    Most existing approaches for image and video compression perform transform coding in the pixel space to reduce redundancy. However, due to the misalignment between the pixel-space distortion and human perception, such schemes often face the difficulties in achieving both high-realism and high-fidelity at ultra-low bitrate. To solve this problem, we propose \

  74. Jun Rao, Xuebo Liu, Hexuan Deng, Zepeng Lin

    In mathematical reasoning, data selection strategies predominantly rely on static, externally defined metrics, which fail to adapt to the evolving capabilities of models during training. This misalignment limits the efficiency of Supervised Fine-Tuning and Reinforcement Learning. To bridge this gap, we introduce SAI-DPO (Self-Aware Iterative Data Persistent

  75. Benjamin Schneider, Dongfu Jiang, Chao Du, Tianyu Pang

    Long-video understanding has emerged as a crucial capability in real-world applications such as video surveillance, meeting summarization, educational lecture analysis, and sports broadcasting. However, it remains computationally prohibitive for VideoLLMs, primarily due to two bottlenecks: 1) sequential video decoding, the process of converting the raw bit s

  76. Ping Liu, Chi Zhang

    To what extent does concept erasure eliminate generative capacity in diffusion models? While prior evaluations have primarily focused on measuring concept suppression under specific textual prompts, we explore a complementary and fundamental question: do current concept erasure techniques genuinely remove the ability to generate targeted concepts, or do they

  77. Liyu Zhong, Sheng Mao

    We derive an exact contrast-expansion formalism for the effective conductivity of heterogeneous materials (media) with local properties described by arbitrary continuous random fields, significantly generalizing the widely used binary-field models. The theory produces a rapidly convergent Neumann-series that, upon Gaussian closure via a Hermite expansion, yi

  78. Abhay Kumara Sri Krishna Nandiraju, Gondy Leroy, David Kauchak, Arif Ahmed

    Understanding health information is essential in achieving and maintaining a healthy life. We focus on simplifying health information for better understanding. With the availability of generative AI, the simplification process has become efficient and of reasonable quality, however, the algorithms remove information that may be crucial for comprehension. In

  79. Mai Lee Chang, Kim Baraka, Greg Trafton, Zach Lalu Vazhekatt

    When agents interact with people as part of a team, fairness becomes an important factor. Prior work has proposed fairness metrics based on teammates' capabilities for task allocation within human-agent teams. However, most metrics only consider teammate capabilities from a third-person point of view (POV). In this work, we extend these metrics to include ta

  80. Hongfei Xue, Yufeng Tang, Jun Zhang, Xuelong Geng

    Although multilingual automatic speech recognition (ASR) systems have significantly advanced, enabling a single model to handle multiple languages, inherent linguistic differences and data imbalances challenge SOTA performance across all languages. While language identification (LID) models can route speech to the appropriate ASR model, they incur high costs

  81. Xiao Hu, Yang Ye

    Robotic manipulation in industrial scenarios such as construction commonly faces uncertain observations in which the state of the manipulating object may not be accurately captured due to occlusions and partial observables. For example, object status estimation during pipe assembly, rebar installation, and electrical installation can be impacted by observati

  82. Yuhao Xue, Zhifei Zhang, Xinyang Jiang, Yifei Shen

    Adversarial attacks exploiting unrestricted natural perturbations present severe security risks to deep learning systems, yet their transferability across models remains limited due to distribution mismatches between generated adversarial features and real-world data. While recent works utilize pre-trained diffusion models as adversarial priors, they still e

  83. Yechan Park, Gyuhyeon Pak, Euntai Kim

    While most people associate LiDAR primarily with its ability to measure distances and provide geometric information about the environment (via point clouds), LiDAR also captures additional data, including reflectivity or intensity values. Unfortunately, when LiDAR is applied to Place Recognition (PR) in mobile robotics, most previous works on LiDAR-based PR

  84. Tianlai Yang, Mo Xiong, Ming Xue, Xinwei Li

    Integer factorization remains a significant challenge for classical computers and is fundamental to the security of RSA encryption. Adiabatic quantum algorithms present a promising solution, yet their practical implementation is limited by the short coherence times of current NISQ devices and quantum simulators. In this work, we apply the chopped random-basi

  85. Mingbo Song, Heming Xia, Jun Zhang, Chak Tou Leong

    Speculative Decoding (SD) has emerged as a widely used paradigm to accelerate the inference of large language models (LLMs) without compromising generation quality. It works by efficiently drafting multiple tokens using a compact model and then verifying them in parallel using the target LLM. Notably, Self-Speculative Decoding proposes skipping certain layer

  86. Liyan Wang, Weixiang Zhou, Cong Wang, Kin-Man Lam

    Ultra-high-definition (UHD) image restoration aims to specifically solve the problem of quality degradation in ultra-high-resolution images. Recent advancements in this field are predominantly driven by deep learning-based innovations, including enhancements in dataset construction, network architecture, sampling strategies, prior knowledge integration, and

  87. Bin Xu, Yu Bai, Huashan Sun, Yiguan Lin

    As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational scenarios, incorporating synthetic data containing 9 major scenarios and over 4,000 distinct educational contexts. To enable c

  88. Tanqiu Jiang, Jiacheng Liang, Rongyi Zhu, Jiawei Zhou

    Large vision-language models (VLMs) are highly vulnerable to multimodal jailbreak attacks that exploit visual-textual interactions to bypass safety guardrails. In this paper, we present DTR, a novel inference-time defense that mitigates multimodal jailbreak attacks through optimizing the model's key-value (KV) caches. Rather than relying on curated safety-sp

  89. Chongjie Si, Yidan Cui, Fuchao Yang, Xiaokang Yang

    Learning from inaccurate annotations has gained significant attention due to the high cost of precise labeling. However, despite the presence of erroneous labels, models trained on noisy data often retain the ability to make accurate predictions. This intriguing phenomenon raises a fundamental yet largely unexplored question: why models can still extract cor

  90. Yiping Shu, Shen Li

    Accurate redshift determinations of both lenses and sources are critical for confirming strong-lens systems and fully realizing their scientific value. However, the thousands of strong-lens candidates now routinely discovered in wide-field imaging surveys make one-by-one follow-up observations impractical. In this work, we investigate the capability and effi

  91. Siu Lun Chau, Michele Caprio, Krikamol Muandet

    Quantifying differences between probability distributions is fundamental to statistics and machine learning, primarily for comparing statistical uncertainty. In contrast, epistemic uncertainty -- due to incomplete knowledge -- requires richer representations than those offered by classical probability. Imprecise probability (IP) theory offers such models, ca

  92. Rui Zhang, Na Zhang, Yapeng Zeng, Tao Yang

    In this paper, Ore extensions of multiplier Hopf coquasigroups are studied. Necessary and sufficient conditions for the Ore extension of a regular multiplier Hopf coquasigroup to be a multiplier Hopf coquasigroup are given. Furthermore, the isomorphism between two such Ore extensions is discussed.

  93. Ji Guo, Long Zhou, Zhijin Wang, Jiaming He

    In recent years, deep learning-based Monocular Depth Estimation (MDE) models have been widely applied in fields such as autonomous driving and robotics. However, their vulnerability to backdoor attacks remains unexplored. To fill the gap in this area, we conduct a comprehensive investigation of backdoor attacks against MDE models. Typically, existing backdoo

  94. B. Bahr-Kalus, D. Parkinson, K. Lodha, E. Mueller

    The peak of the matter power spectrum, known as the turnover (TO) scale, is determined by the horizon size at the time of matter-radiation equality. This scale can serve as a standard ruler, independent of other features in the matter power spectrum, such as baryon acoustic oscillations (BAO). Here, we present the first detection of the turnover in the galax

  95. Bolin Chen, Shanzhi Yin, Hanwei Zhu, Lingyu Zhu

    In this paper, we propose to compress human body video with interactive semantics, which can facilitate video coding to be interactive and controllable by manipulating semantic-level representations embedded in the coded bitstream. In particular, the proposed encoder employs a 3D human model to disentangle nonlinear dynamics and complex motion of human body

  96. Hongchen Wei, Zhenzhong Chen

    Recent advances in Reasoning LLMs (e.g., DeepSeek-R1 and OpenAI-o1) have showcased impressive reasoning capabilities via reinforcement learning. However, extending these capabilities to Multimodal LLMs (MLLMs) is hampered by the prohibitive costs of retraining and the scarcity of high-quality, verifiable multimodal reasoning datasets. This paper introduces F

  97. Cristian Lenart, Satoshi Naito, Daisuke Sagaki, Leonardo C. Mihalcea

    We prove an identity for (torus-equivariant) 3-point, genus 0, $K$-theoretic Gromov-Witten invariants of flag manifolds $G/P$, which can be thought of as a replacement for the ``divisor axiom'' in their (torus-equivariant) quantum $K$-theory. This identity enables us to compute these invariants when two insertions are Schubert classes and the other a Schuber

  98. Zirui Pang, Haosheng Tan, Yuhan Pu, Zhijie Deng

    Image classification benchmark datasets such as CIFAR, MNIST, and ImageNet serve as critical tools for model evaluation. However, despite the cleaning efforts, these datasets still suffer from pervasive noisy labels and often contain missing labels due to the co-existing image pattern where multiple classes appear in an image sample. This results in misleadi

  99. Chongjie Si, Kangtao Lv, Jingjing Jiang, Yadao Wang

    Model merging offers a training-free alternative to multi-task learning by combining independently fine-tuned models into a unified one without access to raw data. However, existing approaches often rely on heuristics to determine the merging coefficients, limiting their scalability and generality. In this work, we revisit model merging through the lens of l

  100. Le Ma, Shirao Yang, Zihao Wang, Yinggui Wang

    The proliferation of large models has intensified the need for efficient data valuation methods to quantify the contribution of individual data providers. Traditional approaches, such as game-theory-based Shapley value and influence-function-based techniques, face prohibitive computational costs or require access to full data and model training details, maki