Skip to content

May 2025 arXiv papers — page 102

Showing 10,10110,200 of 24,552 papers

  1. Zhixin Ma, Chong-Wah Ngo

    Known-item search (KIS) involves only a single search target, making relevance feedback-typically a powerful technique for efficiently identifying multiple positive examples to infer user intent-inapplicable. PicHunter addresses this issue by asking users to select the top-k most similar examples to the unique search target from a displayed set. Under ideal

  2. Amit Mondal, Biswajit Pandey, Anindita Nandi

    The role of large-scale environment in shaping the structural and kinematic properties of stellar halos remains an open question. We investigate whether the cosmic web environments affect the spatial and velocity anisotropies of stellar halos in Milky Way-mass galaxies. Using high-resolution data from the TNG50 simulation, we analyze 29 stellar halos from ea

  3. Tieshuai Song, Jiandong Ye, Ao Guo, Guidong He

    Multi-sensor fusion has significant potential in perception tasks for both indoor and outdoor environments. Especially under challenging conditions such as adverse weather and low-light environments, the combined use of millimeter-wave radar and RGB-D sensors has shown distinct advantages. However, existing multi-sensor datasets in the fields of autonomous d

  4. Xiaozhao Liu, Dinggang Shen, Xihui Liu

    Pretrained generative models have opened new frontiers in brain decoding by enabling the synthesis of realistic texts and images from non-invasive brain recordings. However, the reliability of such outputs remains questionable--whether they truly reflect semantic activation in the brain, or are merely hallucinated by the powerful generative models. In this p

  5. Makram Hamouda, Mohamed Majdoub, Tarek Saanouni

    This work explores the global existence and scattering behavior of solutions to a damped, inhomogeneous nonlinear Schrodinger equation featuring a time-dependent damping term, an inverse-square potential, and an inhomogeneous nonlinearity. We establish global well-posedness in the energy space for subcritical, mass-critical, and energy-critical regimes, usin

  6. Yanshu Li, Jianjiang Yang, Tian Yun, Pinyuan Feng

    Multimodal in-context learning (ICL) has emerged as a key mechanism for harnessing the capabilities of large vision-language models (LVLMs). However, its effectiveness remains highly sensitive to the quality of input ICL sequences, particularly for tasks involving complex reasoning or open-ended generation. A major limitation is our limited understanding of

  7. Shogo Tomizuka, Hajime Kobayashi, Naritaka Oshita, Kazufumi Takahashi

    We study the dynamics of odd-parity perturbations on a static and spherically symmetric black hole background with a timelike vector field based on the effective field theory (EFT) approach. We derive the quadratic Lagrangian written in terms of two master variables, corresponding to the tensor and vector gravitons, which are coupled in general, while they c

  8. Taobo Liao, Taoran Li, Prathamesh Nadkarni

    In this survey, we will explore the interaction between secure multiparty computation and the area of machine learning. Recent advances in secure multiparty computation (MPC) have significantly improved its applicability in the realm of machine learning (ML), offering robust solutions for privacy-preserving collaborative learning. This review explores key co

  9. Ta Duc Huy, Duy Anh Huynh, Yutong Xie, Yuankai Qi

    Visual grounding (VG) is the capability to identify the specific regions in an image associated with a particular text description. In medical imaging, VG enhances interpretability by highlighting relevant pathological features corresponding to textual descriptions, improving model transparency and trustworthiness for wider adoption of deep learning models i

  10. Amitash Nanda, Md Kamal Hossain Chowdhury, Hannah Ross, Kevin Gott

    Load balancing is critical for successful large-scale high-performance computing (HPC) simulations. With modern supercomputers increasing in complexity and variability, dynamic load balancing is becoming more critical to use computational resources efficiently. In this study, performed during a summer collaboration at Lawrence Berkeley National Laboratory, w

  11. Mahmut Yurt, Xin Ye, Yunsheng Ma, Jingru Luo

    3D perception plays an essential role for improving the safety and performance of autonomous driving. Yet, existing models trained on real-world datasets, which naturally exhibit long-tail distributions, tend to underperform on rare and safety-critical, vulnerable classes, such as pedestrians and cyclists. Existing studies on reweighting and resampling techn

  12. Joseph S. Andrews, Andrey I. Bondarev, Per Jönsson, Jon Grumer

    Large-scale multiconfigurational calculations are conducted on experimentally significant transitions in Lr I and its lanthanide homologue Lu I, exhibiting good agreement with recent theoretical and experimental results. A single reference calculation is performed, allowing for substitutions from the core within a sufficiently large active set to effectively

  13. Muniba Noreen, Furqan Shaukat

    Lung cancer remains among the deadliest types of cancer in recent decades, and early lung nodule detection is crucial for improving patient outcomes. The limited availability of annotated medical imaging data remains a bottleneck in developing accurate computer-aided diagnosis (CAD) systems. Self-supervised learning can help leverage large amounts of unlabel

  14. Hongbo Xia, Kaiqiang Yu, Shengxin Liu, Cheng Long

    Cohesive subgraph mining is a fundamental problem in graph theory with numerous real-world applications, such as social network analysis and protein-protein interaction modeling. Among various cohesive subgraphs, the $\gamma$-quasi-clique is widely studied for its flexibility in requiring each vertex to connect to at least a $\gamma$ proportion of other vert

  15. Bowen Jin, Jinsung Yoon, Priyanka Kargupta, Sercan O. Arik

    Reinforcement learning (RL) has demonstrated strong potential in training large language models (LLMs) capable of complex reasoning for real-world problem solving. More recently, RL has been leveraged to create sophisticated LLM-based search agents that adeptly combine reasoning with search engine use. While the use of RL for training search agents is promis

  16. Zehong Wang, Zheyuan Liu, Tianyi Ma, Jiazheng Li

    Graph-structured data pervades domains such as social networks, biological systems, knowledge graphs, and recommender systems. While foundation models have transformed natural language processing, vision, and multimodal learning through large-scale pretraining and generalization, extending these capabilities to graphs -- characterized by non-Euclidean struct

  17. Yifei Chen, Jihao Liu, Yanze Wang

    By applying the theory of the minimal model program for adjoint foliated structures, we establish the Sarkisov program for algebraically integrable foliations on klt varieties: any two Mori fiber spaces of such structure are connected by a sequence of Sarkisov links. Combining with a result of R. Mascharak, we establish the Sarkisov program for foliations in

  18. Han Long, Bingsheng He, Yinyu Ye, Jiheng Zhang

    In this paper, we introduce the Adaptive Inertial Method (AIM), a novel framework for accelerated first-order methods through a customizable inertial term. We provide a rigorous convergence analysis establishing a global convergence rate of O(1/k) under mild conditions, requiring only convexity and local Lipschitz differentiability of the objective function.

  19. Chengkun Zhang, Guangtai Lu, Nattujuks Pholsen, Yasutomo Ota

    We experimentally realized wide-mode-area slow-light modes in valley photonic crystals (VPhCs) heterostructure waveguides. The waveguides are fabricated on a silicon slab by inserting gapless photonic graphene layers with varying widths and modifying the unit cell spacing near the domain walls. By reducing the spacing between unit cells at the domain boundar

  20. Bartłomiej Wróblewski, Gioele Gottardo, Anastasios Zouzias

    We design and implement parallel prefix sum (scan) algorithms using Ascend AI accelerators. Ascend accelerators feature specialized computing units: the cube units for efficient matrix multiplication and the vector units for optimized vector operations. A key feature of the proposed scan algorithms is their extensive use of matrix multiplications and accumul

  21. Ke Guo, Haochen Liu, Xiaojun Wu, Jia Pan

    End-to-end (E2E) autonomous driving systems offer a promising alternative to traditional modular pipelines by reducing information loss and error accumulation, with significant potential to enhance both mobility and safety. However, most existing E2E approaches directly generate plans based on dense bird's-eye view (BEV) grid features, leading to inefficienc

  22. Xuanliang Zhang, Dingzirui Wang, Keyan Xu, Qingfu Zhu

    The table reasoning task, crucial for efficient data acquisition, aims to answer questions based on the given table. Recently, reasoning large language models (RLLMs) with Long Chain-of-Thought (Long CoT) significantly enhance reasoning capabilities, leading to brilliant performance on table reasoning. However, Long CoT suffers from high cost for training an

  23. Chenliang Zhou, Heejin Ahn, Ian M. Mitchell

    In formal safety verification, many proposed algorithms use parametric set representations and convert the computation of the relevant sets into an optimization problem; consequently, the choice of parameterization and objective function have a significant impact on the efficiency and accuracy of the resulting computation. In particular, recent papers have e

  24. Ian Steenstra, Timothy W. Bickmore

    The proliferation of Large Language Models (LLMs) and Intelligent Virtual Agents acting as psychotherapists presents significant opportunities for expanding mental healthcare access. However, their deployment has also been linked to serious adverse outcomes, including user harm and suicide, facilitated by a lack of standardized evaluation methodologies capab

  25. Ziliang Wang, Xuhui Zheng, Kang An, Cijun Ouyang

    Efficient multi-hop reasoning requires Large Language Models (LLMs) based agents to acquire high-value external knowledge iteratively. Previous work has explored reinforcement learning (RL) to train LLMs to perform search-based document retrieval, achieving notable improvements in QA performance, but underperform on complex, multi-hop QA resulting from the s

  26. Mingchao Jiang, Abhinav Jain, Sophia Zorek, Chris Jermaine

    We introduce SIMCOPILOT, a benchmark that simulates the role of large language models (LLMs) as interactive, "copilot"-style coding assistants. Targeting both completion (finishing incomplete methods or code blocks) and infill tasks (filling missing segments within existing code), SIMCOPILOT provides a comprehensive framework for evaluating LLM coding capabi

  27. Fernando Artaza-Covarrubias, Tonatiuh Sánchez-Vizuet, Manuel Solano

    We propose and analyze an HDG scheme for the Laplace-domain interaction between a transient acoustic wave and a bounded elastic solid embedded in an unbounded fluid medium. Two mixed variables (the stress tensor and the velocity of the acoustic wave) are included while the symmetry of the stress tensor is imposed weakly by considering the antisymmetric part

  28. Xiaobao Huang, Yihong Ma, Anjali Gurajapu, Jules Schleinitz

    Reaction virtual screening and discovery are fundamental challenges in chemistry and materials science, where traditional graph neural networks (GNNs) struggle to model multi-reactant interactions. In this work, we propose ChemHGNN, a hypergraph neural network (HGNN) framework that effectively captures high-order relationships in reaction networks. Unlike GN

  29. Aryaman Arora, Neil Rathi, Nikil Roashan Selvam, Róbert Csordás

    State space models (SSMs) for language modelling promise an efficient and performant alternative to quadratic-attention Transformers, yet show variable performance on recalling basic information from the context. While performance on synthetic tasks like Associative Recall (AR) can point to this deficiency, behavioural metrics provide little information as t

  30. Darukeesan Pakiyarajah, Eduardo Pavez, Antonio Ortega, Debargha Mukherjee

    Data-dependent transforms are increasingly being incorporated into next-generation video coding systems such as AVM, a codec under development by the Alliance for Open Media (AOM), and VVC. To circumvent the computational complexities associated with implementing non-separable data-dependent transforms, combinations of separable primary transforms and non-se

  31. Zihu Wang, Boxun Xu, Hejia Geng, Peng Li

    Graph contrastive learning (GCL) has demonstrated great promise for learning generalizable graph representations from unlabeled data. However, conventional GCL approaches face two critical limitations: (1) the restricted expressive capacity of multilayer perceptron (MLP) based encoders, and (2) suboptimal negative samples that either from random augmentation

  32. Frederick Matsuda, Shugo Oguri, Yutaro Sekimoto, Aritoki Suzuki

    LiteBIRD is a JAXA-led international project aimed at measuring the cosmic microwave background (CMB) polarization with high sensitivity to detect polarization $B$ modes. This detection would provide evidence of inflation. LiteBIRD will observe the full sky for three years at the L2 Lagrange point of the Earth-Sun system across 34-448 GHz, and is expected to

  33. Eray Can Elumar, Cem Tekin, Osman Yagan

    Recent advances in large language models (LLMs) have enabled automated dataset labeling with minimal human supervision. While majority voting across multiple LLMs can improve label reliability by mitigating individual model biases, it incurs high computational costs due to repeated querying. In this work, we propose a novel online framework, Cost-aware Major

  34. Zhentao He, Chao Ji

    In this paper, we study the following nonlinear Dirac equations (NLDE) on noncompact metric graph $\mathcal{G}$ with localized nonlinearities \begin{equation} \mathcal{D} u - \omega u= a\chi_{\mathcal{K}}|u|^{p-2}u, \end{equation} where $\mathcal{D}$ is the Dirac operator on $\mathcal{G}$, $u: \mathcal{G} \to \mathbb{C}^2$, $\omega\in \mathbb{R}$, $a > 0$, $

  35. Feifei Shi, Xueyan Yin, Kang Wang, Wanyu Tu

    Time series analysis is pivotal in domains like financial forecasting and biomedical monitoring, yet traditional methods are constrained by limited nonlinear feature representation and long-term dependency capture. The emergence of Large Language Models (LLMs) offers transformative potential by leveraging their cross-modal knowledge integration and inherent

  36. Yihang Li, Tianle Zhang, Xuelong Wei, Jiayi Li

    Robot manipulation learning from human demonstrations offers a rapid means to acquire skills but often lacks generalization across diverse scenes and object placements. This limitation hinders real-world applications, particularly in complex tasks requiring dexterous manipulation. Vision-Language-Action (VLA) paradigm leverages large-scale data to enhance ge

  37. Shuai Zu, Wanqiang Zhu, Fuli Zhang, Chi Xiao

    This paper presents a classification of generator excitation waveforms using principal component analysis (PCA) and machine learning models, including logistic regression, random forest, and gradient boosting decision trees (GBDT). Building upon the traditional Steinmetz equation, a temperature correction term is introduced. Through nonlinear regression and

  38. Samantha Hergott, Viqar Husain, Saeed Rastgoo

    We present an asymptotically flat spherically symmetric non-singular metric that describes gravitational collapse and matter bounce with transient black hole and white hole regions. The metric provides a dynamical counterpart to proposed static non-singular black holes, and a phenomenological model for possible black hole to white hole transitions in quantum

  39. Ishmanbir Singh, Dipankar Srirag, Aditya Joshi

    Sarcasm is a challenge to sentiment analysis because of the incongruity between stated and implied sentiment. The challenge is exacerbated when the implication may be relevant to a specific country or geographical region. Pragmatic metacognitive prompting (PMP) is a cognition-inspired technique that has been used for pragmatic reasoning. In this paper, we ha

  40. Jing Yu, Yuqi Tang, Kehua Feng, Mingyang Rao

    Large Language Models (LLMs) have shown impressive capabilities in contextual understanding and reasoning. However, evaluating their performance across diverse scientific domains remains underexplored, as existing benchmarks primarily focus on general domains and fail to capture the intricate complexity of scientific data. To bridge this gap, we construct Sc

  41. Tianyi Ma, Yiyue Qian, Zheyuan Zhang, Zehong Wang

    The exponential growth of data-driven systems and AI technologies has intensified the demand for high-quality web-sourced datasets. While existing datasets have proven valuable, conventional web data collection approaches face significant limitations in terms of human effort and scalability. Current data-collecting solutions fall into two categories: wrapper

  42. Jason Yang, Wenda Chu, Daniel Khalil, Raul Astudillo

    Protein fitness optimization involves finding a protein sequence that maximizes desired quantitative properties in a combinatorially large design space of possible sequences. Recent advances in steering protein generative models (e.g., diffusion models and language models) with labeled data offer a promising approach. However, most previous studies have opti

  43. Yifeng Meng, Kui Wang

    We prove the existence and uniqueness of the Robin heat kernel on compact Riemannian manifolds with smooth boundary for Robin parameter $\alpha\in\mathbb{R}$, expressed as a spectral expansion in terms of Robin eigenvalues and eigenfunctions. For the non-negative parameter regime ($\alpha\ge 0$), we present a direct proof based on trace Sobolev inequalities

  44. Kangli Wang, Shihao Li, Qianxi Yi, Wei Gao

    Recently, immersive media and autonomous driving applications have significantly advanced through 3D Gaussian Splatting (3DGS), which offers high-fidelity rendering and computational efficiency. Despite these advantages, 3DGS as a display-oriented representation requires substantial storage due to its numerous Gaussian attributes. Current compression methods

  45. Yanshu Li, Jianjiang Yang, Ziteng Yang, Bozheng Li

    Multimodal in-context learning (ICL) is becoming a key capability that allows large vision-language models (LVLMs) to adapt to novel tasks without parameter updates, which expands their usefulness in many real-world applications. However, ICL performance remains unstable even when the in-context demonstrations (ICDs) are well matched, showing that LVLMs stil

  46. Qihang Yu, Kairui Fu, Zheqi Lv, Shengyu Zhang

    Recent advances in large language models (LLMs) have enabled more semantic-aware recommendations through natural language generation. Existing LLM for recommendation (LLM4Rec) methods mostly operate in a System 1-like manner, relying on superficial features to match similar items based on click history, rather than reasoning through deeper behavioral logic.

  47. Dennis Hong, Yusuke Tanaka

    BALLU, the Buoyancy Assisted Lightweight Legged Unit, is a unique legged robot with a helium balloon body and articulated legs \fig{fig:fig1}. Since it is buoyant-assisted, BALLU is inherently stable, never falling over, while being able to walk, jump, and interact safely with people. The BALLU art installation builds on this playful platform to express flui

  48. Sona Elza Simon, Preethi Jyothi

    Effective cross-lingual transfer remains a critical challenge in scaling the benefits of large language models from high-resource to low-resource languages. Towards this goal, prior studies have explored many approaches to combine task knowledge from task-specific data in a (high-resource) source language and language knowledge from unlabeled text in a (low-

  49. Yusuke Masubuchi, Takefumi Hiraki, Yuichi Hiroi, Masanori Ibara

    The digital transformation of smart cities and workplaces requires effective integration of physical and cyber spaces, yet existing digital twin solutions remain limited in supporting real-time, multi-user collaboration. While metaverse platforms enable shared virtual experiences, they have not supported comprehensive integration of IoT sensors on physical s

  50. Yuxuan Wang, Jingshu Chen, Qingyang Wang

    Command injection vulnerabilities are a significant security threat in dynamic languages like Python, particularly in widely used open-source projects where security issues can have extensive impact. With the proven effectiveness of Large Language Models(LLMs) in code-related tasks, such as testing, researchers have explored their potential for vulnerabiliti

  51. Zhiyu Shen, Jiyuan Liu, Yunhe Pang, Yanghui Rao

    Multi-Hop Question Answering (MHQA) is crucial for evaluating the model's capability to integrate information from diverse sources. However, creating extensive and high-quality MHQA datasets is challenging: (i) manual annotation is expensive, and (ii) current synthesis methods often produce simplistic questions or require extensive manual guidance. This pape

  52. V. Rajeswari, Nalin Kant Mohanty

    This topology can achieve a high step-up gain by utilizing a switched capacitor and switched inductor-based VMC network arrangement.Furthermore, the proposed topology can achieve an output gain of approximately three times at a nominal duty ratio with reduced voltage and current stress across the switch, and enhance the maximum efficiency to 96.7

  53. Emma Sulaver

    We develop a framework for factorizing embeddings of non-commutative Sobolev spaces on quantum tori through newly defined Orlicz-Schatten sequence ideals. After introducing appropriate non-commutative Sobolev norms and Orlicz spectral conditions, we establish a summing operator characterization of the quantum Laplacian embedding. Our main results provide bot

  54. Rafael Bravo, Walter Riquelme

    We compute the cross-correlation between the anisotropies of the cosmological gravitational wave background (CGWB) and the galaxy density contrast. We show that the cross-correlation is non-zero due to the {\it late} integrated Sachs-Wolfe (ISW) effect experienced by tensor modes. We study the detection prospects of the cross-correlation signal against cosmi

  55. Jeremy Qin

    Time series forecasting plays a crucial role in various applications, particularly in healthcare, where accurate predictions of future health trajectories can significantly impact clinical decision-making. Ensuring transparency and explainability of the models responsible for these tasks is essential for their adoption in critical settings. Recent work has e

  56. Jia-Mian Li, Bing-Zhao Li

    Interrupted sampling repeater jamming (ISRJ) poses a serious threat to radar target detection. Traditional time-frequency (TF) domain anti-jamming methods are prone to TF aliasing in multi-component signal scenarios, and cannot effectively suppress ISRJ with energy close to the real target under low signal-to-noise ratio (SNR) conditions. To address these ch

  57. Lixi Rao, Jiajun Wang, Xinhao Wang, Shunben Wu

    Topological spin textures, such as merons and skyrmions, have shown significance in both fundamental science and practical applications across diverse physical systems. The optical skyrmionic textures in real space have been extensively explored, but which in momentum space are still rarely studied. Here, we report the experimental generation of momentum-spa

  58. Sirui Li, Linkai Peng, Zheyuan Zhang, Gorkem Durak

    Foundation models (FMs) such as CLIP and SAM have recently shown great promise in image segmentation tasks, yet their adaptation to 3D medical imaging-particularly for pathology detection and segmentation-remains underexplored. A critical challenge arises from the domain gap between natural images and medical volumes: existing FMs, pre-trained on 2D data, st

  59. Sergey Pankov, Georges Harik

    It is straightforward to design an unbiased gradient estimator that stochastically cuts the backpropagation flow through any part of a computational graph. By cutting the parts that have little effect on the computation, one can potentially save a significant amount of backpropagation computation in exchange for a minimal increase in the stochastic gradient

  60. Konstantin M. Dyakonov

    We characterize the Carleson measures $\mu$ on the unit disk for which the image of the Hardy space $H^p$ under the corresponding embedding operator is closed in $L^p(\mu)$. In fact, a more general result involving $(p,q)$-Carleson measures is obtained. A similar problem is solved in the setting of Bergman spaces.

  61. Saehoon Eo, Namhyun Eun, Moon-Jin Kang, HyeonSeop Oh

    In this paper, we study the isothermal gas dynamics. We first establish the global existence of strong solutions to the one-dimensional isothermal Navier-Stokes system for smooth initial data without any smallness conditions, assuming that the initial density has strictly positive lower bound. The existence result allows for possibly degenerate viscosity coe

  62. Alessandro dos Santos Ferreira, Ana Paula Marques Ramos, José Marcato Junior, Wesley Nunes Gonçalves

    Urban forests play a key role in enhancing environmental quality and supporting biodiversity in cities. Mapping and monitoring these green spaces are crucial for urban planning and conservation, yet accurately detecting trees is challenging due to complex landscapes and the variability in image resolution caused by different satellite sensors or UAV flight a

  63. Nanxu Gong, Sixun Dong, Haoyue Bai, Xinyuan Wang

    As a widely-used and practical tool, feature engineering transforms raw data into discriminative features to advance AI model performance. However, existing methods usually apply feature selection and generation separately, failing to strive a balance between reducing redundancy and adding meaningful dimensions. To fill this gap, we propose an agentic featur

  64. Kristine Ann M. Carandang, Jasper Meynard P. Araña, Ethan Robert A. Casin, Christopher P. Monterola

    Due to the legal and ethical responsibilities of healthcare providers (HCPs) for accurate documentation and protection of patient data privacy, the natural variability in the responses of large language models (LLMs) presents challenges for incorporating clinical note generation (CNG) systems, driven by LLMs, into real-world clinical processes. The complexit

  65. Hao Wang, Pinzhi Huang, Jihan Yang, Saining Xie

    The rapid evolution of multimodal large language models (MLLMs) has significantly enhanced their real-world applications. However, achieving consistent performance across languages, especially when integrating cultural knowledge, remains a significant challenge. To better assess this issue, we introduce two new benchmarks: KnowRecall and VisRecall, which eva

  66. Yuhang Zhou, Jing Zhu, Shengyi Qian, Zhuokai Zhao

    Large Language Models (LLMs) are increasingly aligned with human preferences through Reinforcement Learning from Human Feedback (RLHF). Among RLHF methods, Group Relative Policy Optimization (GRPO) has gained attention for its simplicity and strong performance, notably eliminating the need for a learned value function. However, GRPO implicitly assumes a bala

  67. Shu Wang, Jong-Hak Woo, Aaron J. Barth, Vardha N. Bennert

    We present velocity-resolved reverberation lags of H-beta for 20 active galactic nuclei (AGNs) from the Seoul National University AGN Monitoring Project. We detect unambiguous velocity-resolved structures in 12 AGNs, among which eight objects exhibit symmetric structures, two objects show inflow-like characteristics, and two objects display outflow-like sign

  68. Xin Zhou, Weiqing Wang, Francisco J. Baldán, Wray Buntine

    While multimodal data sources are increasingly available from real-world forecasting, most existing research remains on unimodal time series. In this work, we present MoTime, a suite of multimodal time series forecasting datasets that pair temporal signals with external modalities such as text, metadata, and images. Covering diverse domains, MoTime supports

  69. Chen Huang, Junkai Luo, Xinzuo Wang, Wenqiang Lei

    The massive user-generated content (UGC) available in Chinese social media is giving rise to the possibility of studying internet buzzwords. In this paper, we study if large language models (LLMs) can generate accurate definitions for these buzzwords based on UGC as examples. Our work serves a threefold contribution. First, we introduce CHEER, the first data

  70. Cheng Jin, Zhenyu Xiao, Chutao Liu, Yuantao Gu

    Classifier-free guidance (CFG) has emerged as a pivotal advancement in text-to-image latent diffusion models, establishing itself as a cornerstone technique for achieving high-quality image synthesis. However, under high guidance weights, where text-image alignment is significantly enhanced, CFG also leads to pronounced color distortions in the generated ima

  71. Aldo Porco, Dhruv Mehra, Igor Malioutov, Karthik Radhakrishnan

    Learned Sparse Retrieval (LSR) models encode text as weighted term vectors, which need to be sparse to leverage inverted index structures during retrieval. SPLADE, the most popular LSR model, uses FLOPS regularization to encourage vector sparsity during training. However, FLOPS regularization does not ensure sparsity among terms - only within a given query o

  72. Pratik Rakesh Singh, Kritarth Prasad, Mohammadi Zaki, Pankaj Wasnik

    Neural Machine Translation (NMT) systems face significant challenges when working with low-resource languages, particularly in domain adaptation tasks. These difficulties arise due to limited training data and suboptimal model generalization, As a result, selecting an optimal model for translation is crucial for achieving strong performance on in-domain data

  73. Cheng Qian, Hongyi Du, Hongru Wang, Xiusi Chen

    Recent progress in large language models (LLMs) has enabled substantial advances in solving mathematical problems. However, existing benchmarks often fail to reflect the complexity of real-world problems, which demand open-ended, interdisciplinary reasoning and integration of computational tools. To address this gap, we introduce ModelingBench, a novel bench

  74. Simon Chesterman

    "Fake news" is an old problem. In recent years, however, increasing usage of social media as a source of information, the spread of unverified medical advice during the Covid-19 pandemic, and the rise of generative artificial intelligence have seen a rush of legislative proposals seeking to minimize or mitigate the impact of false information spread online.

  75. Koshy George, B. M. Poggianti, B. Vulcani, M. Gullieuszik

    Galaxies undergoing ram-pressure stripping develop gaseous tails that can extend several kiloparsecs outside the galaxy disc. We used far-ultraviolet and H$\alpha$ imaging from the GASP survey to investigate how different stages of stripping affect star formation properties in the tail and disc of 13 galaxies undergoing stripping. These galaxies have differe

  76. Suhas BN, Yash Mahajan, Dominik Mattioli, Andrew M. Sherrill

    This paper investigates the capacity of small language models (0.5B-5B parameters) to generate empathetic responses for individuals with PTSD. We introduce Trauma-Informed Dialogue for Empathy (TIDE), a novel dataset comprising 10,000 two-turn conversations across 500 diverse, clinically-grounded PTSD personas (https://huggingface.co/datasets/yenopoya/TIDE).

  77. Sho Sonoda, Yuka Hashimoto, Isao Ishikawa, Masahiro Ikeda

    Why and when does depth improve generalization? We study this question in an implementation-agnostic state-transition model, where a depth-$k$ predictor is a readout class $H$ composed with the word ball $B(k,F)$ generated by hidden state transitions. Generalization bounds separate implementation error, approximation error, and statistical complexity, and up

  78. Sarfraz Ahmad, Hasan Iqbal, Momina Ahsan, Numaan Naeem

    The rapid adoption of Large Language Models (LLMs) has raised important concerns about the factual reliability of their outputs, particularly in low-resource languages such as Urdu. Existing automated fact-checking systems are predominantly developed for English, leaving a significant gap for the more than 200 million Urdu speakers worldwide. In this work, w

  79. Jiashu He, Jinxuan Fan, Bowen Jiang, Ignacio Hounie

    Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly available. It is essential for solving complex questions in specialized domains where retrieving comprehensive external knowledge is impractical. We propose SAKE (Structured Agentic Knowledge Extrapolation), a RL powered agen

  80. Wen-Chin Huang, Erica Cooper, Tomoki Toda

    We introduce SHEET, a multi-purpose open-source toolkit designed to accelerate subjective speech quality assessment (SSQA) research. SHEET stands for the Speech Human Evaluation Estimation Toolkit, which focuses on data-driven deep neural network-based models trained to predict human-labeled quality scores of speech samples. SHEET provides comprehensive trai

  81. Tirna Deb, Garrett K. Keating, Nikki Zabel, Alessia Moretti

    We present an analysis of the molecular and atomic gas properties of 10 spatially resolved galaxies in the A2626 cluster (z = 0.055), observed as part of the SYMPHANY project. Using CO(2-1) observations from ALMA and SMA, together with HI data from MeerKAT, we examine the interplay between gas phases and environmental influences. A joint morphological and ki

  82. Jhanvi Garg, Krishna Balasubramanian, Quan Zhou

    Simulated tempering is a widely used strategy for sampling from multimodal distributions. In this paper, we consider simulated tempering combined with an arbitrary local Markov chain Monte Carlo sampler and present a new decomposition theorem that provides a lower bound on the restricted spectral gap of the algorithm for sampling from mixture distributions.

  83. Tianbao Zhang, Jian Zhao, Yuer Li, Zheng Zhu

    Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital entertainment, and remote communication. Existing approaches often generate audio-driven facial expressions and gestures

  84. Frederic Wang, Jonathan I. Tamir

    Magnetic Resonance Imaging (MRI) is highly susceptible to motion artifacts due to the extended acquisition times required for k-space sampling. These artifacts can compromise diagnostic utility, particularly for dynamic imaging. We propose a novel alternating minimization framework that leverages a bespoke diffusion model to jointly reconstruct and correct n

  85. Pengfei Huang, Minru Bai

    In this paper, we consider the completely positive tensor decomposition problem with ideal-sparsity. First, we propose an algorithm to generate the maximal cliques of multi-hypergraphs associated with completely positive tensors. This also leads to a necessary condition for tensors to be completely positive. Then, the completely positive tensor decomposition

  86. Hongli Zhou, Hui Huang, Ziqing Zhao, Lvyuan Han

    The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of benchmark effectiveness, examining mainstream prominent LLM ben

  87. Feiyang Cai, Jiahui Bai, Tao Tang, Guijuan He

    Precise recognition, editing, and generation of molecules are essential prerequisites for both chemists and AI systems tackling various chemical tasks. We present MolLangBench, a comprehensive benchmark designed to evaluate fundamental molecule-language interface tasks: language-prompted molecular structure recognition, editing, and generation. To ensure hig

  88. Hemanth Ravipati

    Neuromorphic computing, inspired by the human brain's neural architecture, is revolutionizing artificial intelligence and edge computing with its low-power, adaptive, and event-driven designs. However, these unique characteristics introduce novel cybersecurity risks. This paper proposes Neuromorphic Mimicry Attacks (NMAs), a groundbreaking class of threats t

  89. Tatsuya Abe

    Whereas an extension with non-interference of Hoare logic for sequential programs Owicki--Gries logic ensures the correctness of concurrent programs on strict consistency, it is unsound to weak memory models adopted by modern computer architectures and specifications of programming languages. This paper proposes a novel non-interference notion and provides c

  90. Kevin Hung, Gary Man-Tat Man, Jincheng Wang

    The early detection of Alzheimer's disease (AD) through widespread screening has emerged as a primary strategy to mitigate the significant global impact of AD. EEG measurements offer a promising solution for extensive AD detection. However, the intricate and nonlinear dynamics of multichannel EEG signals pose a considerable challenge for real-time AD diagnos

  91. Haiyang Liu, Yingjie Mao, Xiaoqi Li

    With the rapid development of blockchain technology, various blockchain systems are exhibiting vitality and potential. As a representative of Blockchain 3.0, the EOS blockchain has been regarded as a strong competitor to Ethereum. Nevertheless, compared with Bitcoin and Ethereum, academic research and in-depth analyses of EOS remain scarce. To address this g

  92. Gaurav Kumar, Ayush Garg, Debajyoti Mazumder, Aditya Kishore

    Automated fact-checking has been a challenging task for the research community. Prior work has explored various strategies, such as end-to-end training, retrieval-augmented generation, and prompt engineering, to build robust fact-checking systems. However, their accuracy has not been high enough for real-world deployment. We, on the other hand, propose a new

  93. Kyungho Lee

    Generative UI is transforming interface design by facilitating AI-driven collaborative workflows between designers and computational systems. This study establishes a working definition of Generative UI through a multi-method qualitative approach, integrating insights from a systematic literature review of 127 publications, expert interviews with 18 particip

  94. Yang Luo, Jian-Min Wang

    Accretion disks surrounding supermassive black holes can potentially form stars within the self-gravitating region. These stars undergo high accretion rates because of the dense environment of the active galactic nuclei (AGN) accretion disk. The vorticity of the AGN disk may influence the ultimate mass feeding rate toward the star. In our study, we simulate

  95. Yingming Pu, Tao Lin, Hongyu Chen

    Large Language Model (LLM)-based multi-agent systems (MAS) demonstrate remarkable potential for scientific discovery. Existing approaches, however, often automate scientific discovery using predefined workflows that lack rationality constraints. This often leads to aimless hypothesizing and a failure to consistently link hypotheses with evidence, thereby hin

  96. Yifan Wu, Lutao Yan, Leixian Shen, Yinan Mei

    The emergence of Multi-modal Large Language Models (MLLMs) presents new opportunities for chart understanding. However, due to the fine-grained nature of these tasks, applying MLLMs typically requires large, high-quality datasets for task-specific fine-tuning, leading to high data collection and training costs. To address this, we propose ChartCards, a unifi

  97. Linjie Li, Zhenyu Wu, Yang Ji

    Class-incremental learning (CIL) requires deep learning models to continuously acquire new knowledge from streaming data while preserving previously learned information. Recently, CIL based on pre-trained models (PTMs) has achieved remarkable success. However, prompt-based approaches suffer from prompt overwriting, while adapter-based methods face challenges

  98. Siyue Zhang, Yilun Zhao, Liyuan Geng, Arman Cohan

    Large language model (LLM)-based embedding models, benefiting from large scale pre-training and post-training, have begun to surpass BERT and T5-based models on general-purpose text embedding tasks such as document retrieval. However, a fundamental limitation of LLM embeddings lies in the unidirectional attention used during autoregressive pre-training, whic

  99. Ze Wang, Jingang Qu, Zhenyu Gao, Pascal Morin

    This work demonstrates an airflow inertial based odometry system with multi-sensor data fusion, including thermal anemometer, IMU, ESC, and barometer. This goal is challenging because low-cost IMUs and barometers have significant bias, and anemometer measurements are very susceptible to interference from spinning propellers and ground effects. We employ a GR

  100. Ze Wang, Zhenyu Gao, Jingang Qu, Pascal Morin

    This paper concerns real-time obstacle avoidance for micro aerial vehicles (MAVs). Motivated by teleoperation applications in cluttered environments with limited computational power, we propose a local planner that does not require the knowledge or construction of a global map of the obstacles. The proposed solution consists of a real-time trajectory plannin