Skip to content

December 2024 arXiv papers — page 92

Showing 9,1019,200 of 20,868 papers

  1. Satoshi Eguchi, Makoto Tashiro, Yukikatsu Terada, Hiromitsu Takahashi

    The X-Ray Imaging and Spectroscopy Mission (XRISM) is the 7th Japanese X-ray observatory, whose development and operation are in collaboration with universities and research institutes in Japan, U.S., and Europe, including JAXA, NASA, and ESA. The telemetry data downlinked from the satellite are reduced to scientific products by the pre-pipeline (PPL) and pi

  2. Kaicheng Ni, Heling Su, Jiahui Zhu

    We consider the stochastic incompressible magnetohydrodynamic equations driven by additive jump noises on either the whole space $\mathbb{R}^d$, $d=2,3$ or a smooth bounded domain $D$ in $\mathbb{R}^d$. We establish the local existence and uniqueness of a mild solution in the space $L^q(0,T;\mathbb{L}^{p\otimes}_{\sigma}(D))$ allowing for initial data with l

  3. Saumya Kothari, Harsh Shah, Utkarsh Prajapati, Shrinjay Kaushik

    Mid-cap companies, generally valued between \$2 billion and \$10 billion, provide investors with a well-rounded opportunity between the fluctuation of small-cap stocks and the stability of large-cap stocks. This research builds upon the long-short equity approach (e.g., Michaud, 2018; Dimitriu, Alexander, 2002) customized for mid-cap equities, providing stea

  4. Lanyu Shang, Bozhang Chen, Shiwei Liu, Yang Zhang

    Drought has become a critical global threat with significant societal impact. Existing drought monitoring solutions primarily focus on assessing drought severity using quantitative measurements, overlooking the diverse societal impact of drought from human-centric perspectives. Motivated by the collective intelligence on social media and the computational po

  5. Yasushi Muraki, Shoichi Shibata

    In this paper, the energy levels of the resonant states of toponium, composed of top quark and anti-top quark, are given on the basis of an empirical law. We predict that the mass of the n-th resonant state of toponium is given by Mass(n)=0.81ln}(n) + 347GeV from the empirical law on the resonance level of the bottomonium. The cross-section produced by elect

  6. Emy Mons, Charles Jose

    In hierarchical structure formation, correlations between galaxy properties and their environments reveal important clues about galaxy evolution, emphasizing the importance of measuring these relationships. We probe the environmental dependence of Lyman-break galaxy (LBG) properties in the redshift range of $3$ to $5$ using marked correlation function statis

  7. Zahra Ebrahimi Vargoorani, Ching Yee Suen

    License plate detection (LPD) is essential for traffic management, vehicle tracking, and law enforcement but faces challenges like variable lighting and diverse font types, impacting accuracy. Traditionally reliant on image processing and machine learning, the field is now shifting towards deep learning for its robust performance in various conditions. Curre

  8. Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi

    Recent research arXiv:2410.15027 arXiv:2410.23775 has highlighted the inherent in-context generation capabilities of pretrained diffusion transformers (DiTs), enabling them to seamlessly adapt to diverse visual tasks with minimal or no architectural modifications. These capabilities are unlocked by concatenating self-attention tokens across multiple input an

  9. Vladislav Zubko, Andrey Belyaev

    The primary objective of this paper is to construct an analytical model for determining the total duration of eclipse events of satellites. The approach assumes that the trace formed in the orbital plane, cutting body shadow under the classical conical shadow model, can be described as an ellipse. This allows for the derivation of its parameters through the

  10. Ryo Kishino, Hiroaki Yamagiwa, Ryo Nagata, Sho Yokoi

    Lexical semantic change detection aims to identify shifts in word meanings over time. While existing methods using embeddings from a diachronic corpus pair estimate the degree of change for target words, they offer limited insight into changes at the level of individual usage instances. To address this, we apply Unbalanced Optimal Transport (UOT) to sets of

  11. Noah Schlossberger, Tate McDonald, Kevin Su, Rajavardhan Talashila

    The ability to image electromagnetic fields holds key scientific and industrial applications, including electromagnetic compatibility, diagnostics of high-frequency devices, and experimental scientific work involving field interactions. Generally electric and magnetic field measurements require conductive elements which significantly distort the field. Howev

  12. Seunghee Kim, Changhyeon Kim, Taeuk Kim

    Real-world decision-making often requires integrating and reasoning over information from multiple modalities. While recent multimodal large language models (MLLMs) have shown promise in such tasks, their ability to perform multi-hop reasoning across diverse sources remains insufficiently evaluated. Existing benchmarks, such as MMQA, face challenges due to (

  13. Haonan Xu, Yang Yang

    Out-of-distribution (OOD) detection is crucial for ensuring the reliable deployment of deep models in real-world scenarios. Recently, from the perspective of over-parameterization, a series of methods leveraging weight sparsification techniques have shown promising performance. These methods typically focus on selecting important parameters for in-distributi

  14. Yuhyun Kim, Minwoo Kim, Hyobin Park, Jinwook Jung

    The Multimodal Learning Workshop (PBVS 2024) aims to improve the performance of automatic target recognition (ATR) systems by leveraging both Synthetic Aperture Radar (SAR) data, which is difficult to interpret but remains unaffected by weather conditions and visible light, and Electro-Optical (EO) data for simultaneous learning. The subtask, known as the Mu

  15. Chengyan Wu, Bolei Ma, Zheyu Zhang, Ningyuan Deng

    Aspect-based sentiment analysis (ABSA), a sequence labeling task, has attracted increasing attention in multilingual contexts. While previous research has focused largely on fine-tuning or training models specifically for ABSA, we evaluate large language models (LLMs) under zero-shot conditions to explore their potential to tackle this challenge with minimal

  16. Vaden Masrani, Mohammad Akbari, David Ming Xuan Yue, Ahmad Rezaei

    In the era of costly pre-training of large language models, ensuring the intellectual property rights of model owners, and insuring that said models are responsibly deployed, is becoming increasingly important. To this end, we propose model watermarking via passthrough layers, which are added to existing pre-trained networks and trained using a self-supervis

  17. Zhifei Shi, Zongyao Yin, Sheng Chang, Xiao Yi

    Achieving a balance between computational efficiency and detection accuracy in the realm of rotated bounding box object detection within aerial imagery is a significant challenge. While prior research has aimed at creating lightweight models that enhance computational performance and feature extraction, there remains a gap in the performance of these network

  18. Wenjun Huang, Yang Ni, Hanning Chen, Yirui He

    Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to localize an arbitrary number of targets based on a language expression and continuously track them in a video. This intricate task involves reasoning on multi-modal data and precise target localization with temporal association. However, prior studies overlook the imbalanced

  19. Taeho Hwang, Sukmin Cho, Soyeong Jeong, Hoyun Song

    We introduce EXIT, an extractive context compression framework that enhances both the effectiveness and efficiency of retrieval-augmented generation (RAG) in question answering (QA). Current RAG systems often struggle when retrieval models fail to rank the most relevant documents, leading to the inclusion of more context at the expense of latency and accurac

  20. Xiaowen Chen, Roman Adam, Daniel E. Bürgler, Fangzhou Wang

    Since the discovery of ultrafast demagnetization in Ni thin films in 1996, laser-induced ultrafast spin dynamics have become a prominent research topic in the field of magnetism and spintronics. This development offers new possibilities for the advancement of spintronics and magnetic storage technology. The subject has drawn a substantial number of researche

  21. Xiaotian Xue, Yingdong Xu, Wenjun Ding, Rui Ye

    High-speed photonic integrated circuits leveraging the thin-film lithium niobate (TFLN) platform present a promising approach to address the burgeoning global data traffic demands. As a pivotal component, TFLN-based electro-optic (EO) Mach-Zehnder modulators (MZMs) should exhibit low driving voltage, broad operation bandwidth, high extinction ration, and low

  22. Charles Barthelemy, Ruoyu Chen, Edward Lucyszyn

    Pair trading is a market-neutral quantitative trading strategy that exploits price anomalies between two correlated assets. By taking simultaneous long and short positions, it generates profits based on relative price movements, independent of overall market trends. This study explores the mathematical foundations of pair trading, focusing on identifying coi

  23. G. Saravanakumar, M. Arun

    In this paper, we introduce the concept of $e^\star_{[\gamma,\gamma']}$-open sets in topological spaces and examine their properties in detail. Additionally, we propose a new class of functions, termed $(e^\star_{[\gamma,\gamma']},\ e^\star_{[\beta,\beta']})$-continuous functions, and explore their fundamental characteristics.

  24. Zehua Xia, Yuyang Wu, Yiyun Xia, Cam-Tu Nguyen

    Multi-hop question answering (QA) often requires sequential retrieval (multi-hop retrieval), where each hop retrieves missing knowledge based on information from previous hops. To facilitate more effective retrieval, we aim to distill knowledge from a posterior retrieval, which has access to posterior information like an answer, into a prior retrieval used d

  25. Komal Negi, Ayaka Shimizu, Yoshiro Yaguchi, Madeti Prabhakar

    The set of all virtual or classical braid diagrams forms a monoid and gives a natural monoid action on a direct product of ${\mathbb Z}$ called the up-down action. In this paper, we determine the orbit of every tuple of ${\mathbb Z}$ under the up-down action of virtual or classical braid diagrams. Moreover, we determine the orbit for irreducible braid diagra

  26. Sparsh Pekhale, Rakshith Sathish, Sathisha Basavaraju, Divya Sharma

    Land-use and land cover (LULC) analysis is critical in remote sensing, with wide-ranging applications across diverse fields such as agriculture, utilities, and urban planning. However, automating LULC map generation using machine learning is rendered challenging due to noisy labels. Typically, the ground truths (e.g. ESRI LULC, MapBioMass) have noisy labels

  27. Jari Taskinen

    We study spectra of Toeplitz operators $T_a $ with periodic symbols in Bergman spaces $A^2(\Pi)$ on unbounded periodic planar domains $\Pi$, which are defined as the union of infinitely many copies of the translated, bounded periodic cell $\varpi$. We introduce Floquet-transform techniques and prove a version of the band-gap-spectrum formula, which is well-k

  28. Xinlong Cheng, Tiantian Cao, Guoan Cheng, Bangxuan Huang

    In this work, we address the limitations of denoising diffusion models (DDMs) in image restoration tasks, particularly the shape and color distortions that can compromise image quality. While DDMs have demonstrated a promising performance in many applications such as text-to-image synthesis, their effectiveness in image restoration is often hindered by shape

  29. Shiyan Zhong

    In recent years, a new subclass of tidal disruption events (TDEs) was reported from the literature. The light curve of these TDEs show a re-brightening feature in the decline phase after the first peak, which then leads to a second flare. The re-brightening TDEs challenges the existing light curve fitting tools, which are designed to handle single flare. In

  30. C. Blake, C. Garcia-Quintero, S. Ahlen, D. Bianchi

    The current generation of large galaxy surveys will test the cosmological model by combining multiple types of observational probes. Realising the statistical promise of these new datasets requires rigorous attention to all aspects of analysis including cosmological measurements, modelling, covariance and parameter likelihood. In this paper we present the re

  31. Ziang Wang, Lei Wang, Qi Yi, Yimin Liu

    Unmanned aerial vehicles (UAVs) have played an increasingly important role in military operations and social life. Among all application scenarios, multi-target tracking tasks accomplished by UAV swarms have received extensive attention. However, when UAVs use radar to track targets, the tracking performance can be severely compromised by jammers. To track t

  32. Li Chen, Sheng-Li Qin, Tie Liu, Paul F. Goldsmith

    Interstellar molecules are excellent tools for studying the physical and chemical environments of massive star-forming regions. In particular, vibrationally excited HC$_3$N (HC$_3$N*) lines are the key tracers for probing hot cores environments. We present the Atacama Large Millimeter/submillimeter Array (ALMA) 3 mm observations of HC$_3$N* lines in 60 hot c

  33. Haruka Fukihara, Daisuke Takaishi, Yoshiaki Misugi, Megumi Sasaki

    In this study, we perform 3D magnetohydrodynamics (MHD) simulations of filamentary molecular clouds. We then generate synthetic observations based on the simulation results. Using these, we investigate how the new polarization data analysis method recently introduced by Doi et al. (2021) reflects the magnetic field structure in turbulent filamentary molecula

  34. Hao Wang, Boyi Liu, Yufeng Zhang, Jie Chen

    Competition-level code generation tasks pose significant challenges for current state-of-the-art large language models (LLMs). For example, on the LiveCodeBench-Hard dataset, models such as O1-Mini and O1-Preview achieve pass@1 rates of only 0.366 and 0.143, respectively. While tree search techniques have proven effective in domains like mathematics and gene

  35. Zhen Li, Tan Li, Hai Liu, Tse-Tin Chan

    Proactive caching is essential for minimizing latency and improving Quality of Experience (QoE) in multi-server edge networks. Federated Deep Reinforcement Learning (FDRL) is a promising approach for developing cache policies tailored to dynamic content requests. However, FDRL faces challenges such as an expanding caching action space due to increased conten

  36. Katie Seaborn

    We humans are biased - and our robotic creations are biased, too. Bias is a natural phenomenon that drives our perceptions and behavior, including when it comes to socially expressive robots that have humanlike features. Recognizing that we embed bias, knowingly or not, within the design of such robots is crucial to studying its implications for people in mo

  37. Tao Fang, Derek F. Wong, Lusheng Zhang, Keyan Jin

    While large-scale language models (LLMs) have demonstrated remarkable capabilities in specific natural language processing (NLP) tasks, they may still lack proficiency compared to specialized models in certain domains, such as grammatical error correction (GEC). Drawing inspiration from the concept of curriculum learning, we have delved into refining LLMs in

  38. Austin Cheng, Alston Lo, Kin Long Kelvin Lee, Santiago Miret

    Molecular structure elucidation is a fundamental step in understanding chemical phenomena, with applications in identifying molecules in natural products, lab syntheses, forensic samples, and the interstellar medium. We consider the task of predicting a molecule's all-atom 3D structure given only its molecular formula and moments of inertia, motivated by the

  39. Vidhi Agrawal, Eesha Khalid, Tianyu Tan, Doris Xu

    This study applies machine learning to predict S&P 500 membership changes: key events that profoundly impact investor behavior and market dynamics. Quarterly data from WRDS datasets (2013 onwards) was used, incorporating features such as industry classification, financial data, market data, and corporate governance indicators. Using a Random Forest model, we

  40. Hankun Kang, Jianhao Chen, Yongqi Li, Xin Miao

    Toxicity detection is crucial for maintaining the peace of the society. While existing methods perform well on normal toxic contents or those generated by specific perturbation methods, they are vulnerable to evolving perturbation patterns. However, in real-world scenarios, malicious users tend to create new perturbation patterns for fooling the detectors. F

  41. Deep Bhatt, Surya Ayyagari, Anuruddh Mishra

    Diagnostic errors in healthcare persist as a critical challenge, with increasing numbers of patients turning to online resources for health information. While AI-powered healthcare chatbots show promise, there exists no standardized and scalable framework for evaluating their diagnostic capabilities. This study introduces a scalable benchmarking methodology

  42. Juan P. Mendez, Denis Mamaluy

    We report a nanoscale device concept based on a highly doped $\delta$-layer tunnel junction embedded in a semiconductor for charge sensing. Recent advances in Atomic Precision Advanced Manufacturing (APAM) processes have enabled the fabrication of devices based on quasi-2D, highly conductive, highly doped regions, known as $\delta$-layers, in semiconductor m

  43. Kristijan Kilassa Kvaternik

    For the family of Lozi maps, we study homoclinic points for the saddle fixed point $X$ in the first quadrant. Specifically, in the parameter space, we examine the boundary of the region in which homoclinic points for $X$ exist. For all parameters on that boundary, all intersections of the stable and unstable manifold of $X$, apart from $X$, are tangential, o

  44. Xiaoyang Huang, Andrew Lucas

    We construct a nonlinear fluctuating hydrodynamic effective field theory for Galilean-invariant quantum Hall systems with spontaneously broken translational symmetry. Neglecting the role of energy conservation in a low-temperature regime, the hydrodynamic mode is a magnetophonon with quartic attenuation: $\omega\sim \pm k^2-\mathrm{i} k^z$ with $z=4$. Howeve

  45. Ryan M. T. White, Peter Tuthill

    Wolf-Rayet stars embody the final stable phase of the most massive stars immediately before their evolution is terminated in a supernova explosion. They are responsible for some of the most extreme and energetic phenomena in stellar physics, driving fast and dense stellar winds that are powered by extraordinarily high mass-loss rates arising from their near

  46. Liuke Lyu, Menghan Song, Ting-Tung Wang, Zi Yang Meng

    Entanglement microscopy reveals the true quantum correlations among the microscopic building blocks of many-body systems [Nat. Commun. 16, 96 (2025)]. Using this approach, we study the multipartite entanglement of the quantum Ising model in 1d, 2d, and 3d. We first obtain the full reduced density matrix (tomography) of subregions that have at most 4 sites vi

  47. Iman Khazrak, Shakhnoza Takhirova, Mostafa M. Rezaee, Mehrdad Yadollahi

    The development of accurate medical image classification models is often constrained by privacy concerns and data scarcity for certain conditions, leading to small and imbalanced datasets. To address these limitations, this study explores the use of generative models, such as Denoising Diffusion Probabilistic Models (DDPM) and Progressive Growing Generative

  48. Zhenyu Xiao, Zhe Li, Lipeng Zhu, Boyu Ning

    This paper investigates movable antenna (MA) aided non-orthogonal multiple access (NOMA) for multi-user downlink communication, where the base station (BS) is equipped with a fixed-position antenna (FPA) array to serve multiple MA-enabled users. An optimization problem is formulated to maximize the minimum achievable rate among all the users by jointly optim

  49. Lorenzo Pompili

    We study the Miura map of the KP-II equation on $\mathbb R^2$ and the resulting B\"acklund transform, which adds a line soliton to a given solution. This work aims to develop a complementary approach to T. Mizumachi's method for the $L^2$-stability of the line soliton, which the potential for generalization to multisolitons. We construct the B\"acklund trans

  50. Thomas R. Scruby, Kae Nemoto, Zhenyu Cai

    We show how looped pipeline architectures - which use short-range shuttling of physical qubits to achieve a finite amount of non-local connectivity - can be used to efficiently implement the fault-tolerant non-Clifford gate between 2D surface codes described in (Sci. Adv. 6, eaay4929 (2020)). The shuttling schedule needed to implement this gate is only margi

  51. Amer Abu Arisheh, Jeffrey A. Nanzer

    We present a new approach to secure wireless operations using a simple dipole antenna with a dynamic unbalanced feeding structure. By rapidly switching between two states, a dynamic radiation pattern is generated, resulting in directional modulation. The current distribution on the arms of the dipole antenna are made asymmetric by the balun, which changes th

  52. Hyuhng Joon Kim, Youna Kim, Sang-goo Lee, Taeuk Kim

    Large Language Models (LLMs) demonstrate exceptional performance across diverse tasks by leveraging pre-trained (i.e., parametric) and external (i.e., contextual) knowledge. While substantial efforts have been made to enhance the utilization of both forms of knowledge, situations in which models lack relevant information remain underexplored. To investigate

  53. Giovanni Maria Tomaselli

    Several models of physics beyond the Standard Model predict the existence of new ultralight bosons. This thesis investigates a way to discover such particles through observations of gravitational waves from binary black holes. This is possible through black hole superradiance, which spontaneously creates a "boson cloud" around a rapidly spinning black hole.

  54. Ruixin Mao, Aoyu Shen, Lin Tang, Jun Zhou

    Event-based cameras feature high temporal resolution, wide dynamic range, and low power consumption, which is ideal for high-speed and low-light object detection. Spiking neural networks (SNNs) are promising for event-based object recognition and detection due to their spiking nature but lack efficient training methods, leading to gradient vanishing and high

  55. Shuta Funakoshi, Tatsuo Kobayashi, Hajime Otsuka

    We discuss selection rules of chiral matters in type IIA intersecting and IIB magnetized D-brane models on toroidal orbifolds. Since the chiral matters on toroidal orbifolds are labeled by a certain conjugacy class of the gauged orbifold group, the selection rules involve non-trivial fusion rules. We find that the representation of the chiral matters is desc

  56. Ruihong Zeng, Jinyuan Fang, Siwei Liu, Zaiqiao Meng

    Memory plays a pivotal role in enabling large language model~(LLM)-based agents to engage in complex and long-term interactions, such as question answering (QA) and dialogue systems. While various memory modules have been proposed for these tasks, the impact of different memory structures across tasks remains insufficiently explored. This paper investigates

  57. Umang Bhaskar, Soumyajit Pyne

    The Hotelling-Downs model is a natural and appealing model for understanding strategic positioning by candidates in elections. In this model, voters are distributed on a line, representing their ideological position on an issue. Each candidate then chooses as a strategy a position on the line to maximize her vote share. Each voter votes for the nearest candi

  58. Geling Liu, Yunzhi Tan, Ruichao Zhong, Yuanzhen Xie

    Recently, large language models (LLMs) have significantly improved the performance of text-to-SQL systems. Nevertheless, many state-of-the-art (SOTA) approaches have overlooked the critical aspect of system robustness. Our experiments reveal that while LLM-driven methods excel on standard datasets, their accuracy is notably compromised when faced with advers

  59. Shisong Li, Yongchao Ma, Kang Ma, Weibo Liu

    With the adoption of the revised International System of Units (SI), the Kibble balance has become a pivotal instrument for mass calibrations against the Planck constant, $h$. One of the major focuses in the Kibble balance community is prioritizing experiments that achieve both high accuracy and compactness. The Tsinghua tabletop Kibble balance experiment se

  60. Reza Hadadi

    This paper explores the controllability and state tracking of ensembles from the perspective of optimal transport theory. Ensembles, characterized as collections of systems evolving under the same dynamics but with varying initial conditions, are a fundamental concept in control theory and applications. By leveraging optimal transport, we provide a novel fra

  61. Kan Zheng, Rongtao Xu, Jie Mei, Haojun Yang

    The Ambient Internet of Things (A-IoT) has emerged as a critical direction for achieving effective connectivity as the IoT system evolves to 6G. However, the introduction of A-IoT technologies, particularly involving backscatter modulation, poses numerous challenges for system design and network operations. This paper surveys current standardization efforts,

  62. Uihyeon Jeong, Taegyu Kim

    We consider the Calogero--Moser derivative nonlinear Schr\"odinger equation (CM-DNLS), an $L^2$-critical nonlinear Schr\"odinger type equation enjoying a number of numerous structures, such as nonlocal nonlinearity, self-duality, pseudo-conformal symmetry, and complete integrability. In this paper, we construct smooth finite-time blow-up solutions to (CM-DNL

  63. Hongpeng Cao, Yanbing Mao, Lui Sha, Marco Caccamo

    Real-world accidents in learning-enabled CPS frequently occur in challenging corner cases. During the training of deep reinforcement learning (DRL) policy, the standard setup for training conditions is either fixed at a single initial condition or uniformly sampled from the admissible state space. This setup often overlooks the challenging but safety-critica

  64. Xiaohu Han, Pedro Ribeiro, Stefano Chesi

    We analyze the fate of spiral order in a one-dimensional system of localized magnetic moments coupled to itinerant electrons under a voltage bias. Within an adiabatic approximation for the dynamics of the localized spins, and in the presence of a phenomenological damping term, we demonstrate the occurrence of various dynamical regimes: At small bias a rigidl

  65. Max Mason, Waasi A Jagirdar, David Huang, Rahul Murugan

    The primary objective of this research is to build a Momentum Transformer that is expected to outperform benchmark time-series momentum and mean-reversion trading strategies. We extend the ideas introduced in the paper Trading with the Momentum Transformer: An Intelligent and Interpretable Architecture to equities as the original paper primarily only builds

  66. Peng Gao, Liangyi Zhao

    We establish upper bounds for shifted moments of modular $L$-functions to a fixed modulus as well as quadratic twists of modular $L$-functions under the generalized Riemann hypothesis. Our results are then used to establish bounds for moments of sums involving with Fourier coefficients of a given modular form twisted by Dirichlet characters.

  67. Daniele Agostini, Lakshmi Ramesh, Dawei Shen

    The ABCT variety is defined as the closure of the image of $G(2,n)$ under the Veronese map. We realize the ABCT variety $V(3,n)$ as the determinantal variety of a vector bundle morphism. We use this to give a recursive formula for the fundamental class of $V(3,n)$. As an application, we show that special Schubert coefficients of this class are given by Euler

  68. Rabimba Karanjai, Sam Blackshear, Lei Xu, Weidong Shi

    The growing adoption of formal verification for smart contracts has spurred the development of new verifiable languages like Move. However, the limited availability of training data for these languages hinders effective code generation by large language models (LLMs). This paper presents ConMover, a novel framework that enhances LLM-based code generation for

  69. Yun Liu, Xuechen Liu, Xiaoxiao Miao, Junichi Yamagishi

    Target speaker extraction (TSE) is essential in speech processing applications, particularly in scenarios with complex acoustic environments. Current TSE systems face challenges in limited data diversity and a lack of robustness in real-world conditions, primarily because they are trained on artificially mixed datasets with limited speaker variability and un

  70. Dongjun Hwang, Sungwon Woo, Tom Gao, Raymond Luo

    As Generative AI continues to become more accessible, the case for robust detection of generated images in order to combat misinformation is stronger than ever. Invisible watermarking methods act as identifiers of generated content, embedding image- and latent-space messages that are robust to many forms of perturbations. The majority of current research inv

  71. Bohan Li, Jiannan Guan, Longxu Dou, Yunlong Feng

    The Myers-Briggs Type Indicator (MBTI) is one of the most influential personality theories reflecting individual differences in thinking, feeling, and behaving. MBTI personality detection has garnered considerable research interest and has evolved significantly over the years. However, this task tends to be overly optimistic, as it currently does not align w

  72. Kayla Schroeder, Zach Wood-Doughty

    Large Language Models (LLMs) have become increasingly powerful and ubiquitous, but their stochastic nature poses challenges to the reliability of their outputs. While deterministic settings can improve consistency, they do not guarantee reliability, as a single sample from the model's probability distribution can still be misleading. Building upon the concep

  73. Xiongfeng Zhan, Xueyi Huang

    In combinatorics, P\'{o}lya's Enumeration Theorem is a powerful tool for solving a wide range of counting problems, including the enumeration of groups, graphs, and chemical compounds. In this paper, we present an extension of P\'{o}lya's Enumeration Theorem. As an application, we derive a formula that expresses the $n$-th elementary symmetric polynomial in

  74. Abraham G. Taye, Sador Yemane, Eshetu Negash, Yared Minwuyelet

    In the ever-evolving landscape of medical diagnostics, this study details the systematic design process and concept selection methodology for developing an advanced digital stethoscope, demonstrating the evolution from traditional acoustic models to AI-enhanced digital solutions. The device integrates cutting-edge AI technology with traditional auscultation

  75. Qi Wu, Janick Martinez Esturo, Ashkan Mirzaei, Nicolas Moenne-Loccoz

    3D Gaussian Splatting (3DGS) enables efficient reconstruction and high-fidelity real-time rendering of complex scenes on consumer hardware. However, due to its rasterization-based formulation, 3DGS is constrained to ideal pinhole cameras and lacks support for secondary lighting effects. Recent methods address these limitations by tracing the particles instea

  76. Robert I. Saye

    We design and investigate a variety of multigrid solvers for high-order local discontinuous Galerkin methods applied to elliptic interface and multiphase Stokes problems. Using the template of a standard multigrid V-cycle, we consider a variety of element-wise block smoothers, including Jacobi, multi-coloured Gauss-Seidel, processor-block Gauss-Seidel, and w

  77. Mingxu Chai, Ziyu Shen, Chong Zhang, Yue Zhang

    Document parsing is essential for analyzing complex document structures and extracting fine-grained information, supporting numerous downstream applications. However, existing methods often require integrating multiple independent models to handle various parsing tasks, leading to high complexity and maintenance overhead. To address this, we propose DocFusio

  78. Hong Liu, Saisai Gong, Yixin Ji, Kaixin Wu

    With the rapid advancement of pre-trained large language models (LLMs), recent endeavors have leveraged the capabilities of LLMs in relevance modeling, resulting in enhanced performance. This is usually done through the process of fine-tuning LLMs on specifically annotated datasets to determine the relevance between queries and items. However, there are two

  79. Yakun Niu, Pei Chen, Lei Zhang, Hongjian Yin

    Image Splicing Localization (ISL) is a fundamental yet challenging task in digital forensics. Although current approaches have achieved promising performance, the edge information is insufficiently exploited, resulting in poor integrality and high false alarms. To tackle this problem, we propose a multi-scale cross-fusion and edge-supervision network for ISL

  80. Yan Zhang, Gangyan Zeng, Huawen Shen, Daiqing Wu

    Video text-based visual question answering (Video TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. Inspired by the development of TextVQA in image domain, existing Video TextVQA approaches leverage a language model (e.g. T5) to process text-rich multiple frames and generate answe

  81. Wenbin An, Haonan Lin, Jiahao Nie, Feng Tian

    Generalized Category Discovery is a significant and complex task that aims to identify both known and undefined novel categories from a set of unlabeled data, leveraging another labeled dataset containing only known categories. The primary challenges stem from model bias induced by pre-training on only known categories and the lack of precise supervision for

  82. Sina Bagheri Nezhad, Ameeta Agrawal, Rhitabrat Pokharel

    Multilingual language models (MLLMs) are crucial for handling text across various languages, yet they often show performance disparities due to differences in resource availability and linguistic characteristics. While the impact of pre-train data percentage and model size on performance is well-known, our study reveals additional critical factors that signi

  83. Yingshui Tan, Boren Zheng, Baihui Zheng, Kerui Cao

    With the rapid advancement of Large Language Models (LLMs), significant safety concerns have emerged. Fundamentally, the safety of large language models is closely linked to the accuracy, comprehensiveness, and clarity of their understanding of safety knowledge, particularly in domains such as law, policy and ethics. This factuality ability is crucial in det

  84. Hongbin Zhang, Kehai Chen, Xuefeng Bai, Yang Xiang

    Large language models (LLMs) have exhibited impressive multilingual reasoning capabilities, driven by extensive multilingual pre-training corpora and instruction fine-tuning data. However, a performance gap exists between high- and low-resource language reasoning tasks due to the language imbalance in the pre-training corpus, which is exacerbated by evaluati

  85. Sho Inoue, Kun Zhou, Shuai Wang, Haizhou Li

    Emotional text-to-speech synthesis (TTS) aims to generate realistic emotional speech from input text. However, quantitatively controlling multi-level emotion rendering remains challenging. In this paper, we propose a flow-matching based emotional TTS framework with a novel approach for emotion intensity modeling to facilitate fine-grained control over emotio

  86. Xin Yi, Shunfan Zheng, Linlin Wang, Gerard de Melo

    The emergence of finetuning-as-a-service has revealed a new vulnerability in large language models (LLMs). A mere handful of malicious data uploaded by users can subtly manipulate the finetuning process, resulting in an alignment-broken model. Existing methods to counteract fine-tuning attacks typically require substantial computational resources. Even with

  87. Mingjia Shi, Yuhao Zhou, Ruiji Yu, Zekai Li

    Vision Mamba has shown close to state of the art performance on computer vision tasks, drawing much interest in increasing it's efficiency. A promising approach is token reduction (that has been successfully implemented in ViTs). Pruning informative tokens in Mamba leads to a high loss of key knowledge and degraded performance. An alternative, of merging tok

  88. R. Pablo Arribillaga, Agustin G. Bonifacio

    In the problem of fully allocating an infinitely divisible commodity among agents whose preferences are single-peaked, we show that the uniform rule is the only allocation rule that satisfies efficiency, the equal division guarantee, consistency, and non-obvious manipulability.

  89. Rui Wang, Kaitao Meng, Deshi Li

    Unmanned aerial vehicles (UAVs) have attracted plenty of attention due to their high flexibility and enhanced communication ability. However, the limited coverage and energy of UAVs make it difficult to provide timely wireless service for large-scale sensor networks, which also exist in multiple UAVs. To this end, the advanced collaboration mechanism of UAVs

  90. Jinghan Zeng, Eugene Wu, Sanjay Krishnan

    Many computer systems are now being redesigned to incorporate LLM-powered agents, enabling natural language input and more flexible operations. This paper focuses on handling database transactions created by large language models (LLMs). Transactions generated by LLMs may include semantic errors, requiring systems to treat them as long-lived. This allows for

  91. Qingtao Pan, Wenhao Qiao, Jingjiao Lou, Bing Ji

    Semi-supervised medical image segmentation (SSMIS) uses consistency learning to regularize model training, which alleviates the burden of pixel-wise manual annotations. However, it often suffers from error supervision from low-quality pseudo labels. Vision-Language Model (VLM) has great potential to enhance pseudo labels by introducing text prompt guided mul

  92. Rohit Sehgal, Vishal Tanna, Vinicius Petrucci, Anil Godbole

    High-Performance Computing (HPC) and Artificial Intelligence (AI) workloads typically demand substantial memory bandwidth and, to a degree, memory capacity. CXL memory expansion modules, also known as CXL "type-3" devices, enable enhancements in both memory capacity and bandwidth for server systems by utilizing the CXL protocol which runs over the PCIe inter

  93. Liang-Hong Mo, Zhenyu Xiao, Roderich Moessner, Hongzheng Zhao

    Potential disorder in 1D leads to Anderson localization of the entire spectrum. Upon sacrificing hermiticity by adding non-reciprocal hopping, the non-Hermitian skin effect competes with localization. We find another route for delocalization, which involves imaginary potential disorder. While an entirely random potential generally still leads to localization

  94. Hongyi Jin, Ruihang Lai, Charlie F. Ruan, Yingcheng Wang

    The recent advances in LLMs bring a strong demand for efficient system support to improve overall serving efficiency. As LLM inference scales towards multiple GPUs and even multiple compute nodes, various coordination patterns, such as prefill-decode disaggregation and context migration, arise in serving systems. Most inference services today expose a coarse

  95. Yicheng Feng, Yuetao Chen, Kaiwen Chen, Jingzong Li

    Simulation offers unique values for both enumeration and extrapolation purposes, and is becoming increasingly important for managing the massive machine learning (ML) clusters and large-scale distributed training jobs. In this paper, we build Echo to tackle three key challenges in large-scale training simulation: (1) tracing the runtime training workloads at

  96. Hongjin Qian, Zheng Liu, Peitian Zhang, Zhicheng Dou

    Processing long contexts poses a significant challenge for large language models (LLMs) due to their inherent context-window limitations and the computational burden of extensive key-value (KV) activations, which severely impact efficiency. For information-seeking tasks, full context perception is often unnecessary, as a query's information needs can dynamic

  97. Mingyao Cui, Qunsong Zeng, Kaibin Huang

    Rydberg Atomic REceiver (RARE) is driving a paradigm shift in electromagnetic (EM) wave measurement by harnessing the electron transition phenomenon of Rydberg atoms. Operating at the quantum scale, such receivers have the potential to breakthrough the performance limit of classic receivers, sparking a revolution in physical-layer wireless communications. Th

  98. Samuel Yen-Chi Chen

    Recent advancements in quantum computing (QC) and machine learning (ML) have garnered significant attention, leading to substantial efforts toward the development of quantum machine learning (QML) algorithms to address a variety of complex challenges. The design of high-performance QML models, however, requires expert-level knowledge, posing a significant ba

  99. Shibing Mo, Kai Wu, Qixuan Gao, Xiangyi Teng

    In real-world applications, spectral Graph Neural Networks (GNNs) are powerful tools for processing diverse types of graphs. However, a single GNN often struggles to handle different graph types-such as homogeneous and heterogeneous graphs-simultaneously. This challenge has led to the manual design of GNNs tailored to specific graph types, but these approach

  100. Ritwika Chattopadhyay, Abhishek Malichkar, Zhixuan Ren, Xinyue Zhang

    This paper addresses the challenges faced in large-volume trading, where executing substantial orders can result in significant market impact and slippage. To mitigate these effects, this study proposes a volatility-volume-based order slicing strategy that leverages Exponential Weighted Moving Average and Markov Chain Monte Carlo simulations. These methods a