Skip to content

March 2025 arXiv papers — page 25

Showing 2,4012,500 of 23,633 papers

  1. YangTian Yan, Jinyu Tian

    Deep neural networks (DNNs) are susceptible to Universal Adversarial Perturbations (UAPs), which are instance agnostic perturbations that can deceive a target model across a wide range of samples. Unlike instance-specific adversarial examples, UAPs present a greater challenge as they must generalize across different samples and models. Generating UAPs typica

  2. Yiren Lu, Yunlai Zhou, Yiran Qiao, Chaoda Song

    Open-vocabulary querying in 3D space is crucial for enabling more intelligent perception in applications such as robotics, autonomous systems, and augmented reality. However, most existing methods rely on 2D pixel-level parsing, leading to multi-view inconsistencies and poor 3D object retrieval. Moreover, they are limited to static scenes and struggle with d

  3. Daichi Sasaki, Junna Sugiyama, Kyohei Yamada, Bryce Bixler

    We present the design methodology and characterization of a superconducting magnetic bearing (SMB) system for the polarization modulator in the SAT-LF, one of the small aperture telescopes (SATs) in the Simons Observatory (SO) that is sensitive at 30/40 GHz frequency bands. SO is a ground-based cosmic microwave background (CMB) polarization experiment, with

  4. Ziheng Mao, Yuan He, Jia Zhang, Yimiao Sun

    Heart rate recovery (HRR) within the initial minute following exercise is a widely utilized metric for assessing cardiac autonomic function in individuals and predicting mortality risk in patients with cardiovascular disease. However, prevailing solutions for HRR monitoring typically involve the use of specialized medical equipment or contact wearable sensor

  5. Jaewoo Jeong, Seohee Lee, Daehee Park, Giwon Lee

    Pedestrian trajectory forecasting is crucial in various applications such as autonomous driving and mobile robot navigation. In such applications, camera-based perception enables the extraction of additional modalities (human pose, text) to enhance prediction accuracy. Indeed, we find that textual descriptions play a crucial role in integrating additional mo

  6. Haomin Zhang, Sizhe Shan, Haoyu Wang, Zihao Chen

    Creating high-quality sound effects from videos and text prompts requires precise alignment between visual and audio domains, both semantically and temporally, along with step-by-step guidance for professional audio generation. However, current state-of-the-art video-guided audio generation models often fall short of producing high-quality audio for both gen

  7. Long Gao, Yunhe Zhang, Langkun Chen, Yan Jiang

    Object tracking based on hyperspectral video attracts increasing attention to the rich material and motion information in the hyperspectral videos. The prevailing hyperspectral methods adapt pretrained RGB-based object tracking networks for hyperspectral tasks by fine-tuning the entire network on hyperspectral datasets, which achieves impressive results in c

  8. Kohei Iwaki, Seiya Kato, Shotaro Sakurai

    In this paper, we study the isomonodromy systems associated with the Garnier systems of type 9/2 and type 5/2+3/2. We show that the both of isomonodromy systems admit the singularity reduction (restriction to a movable pole), and the resulting linear differential equations are isomonodromic with respect to the second variable of the Garnier systems. Furtherm

  9. Yang Liu, Xun Zhang, Jiale Du, Xinbo Gao

    Zero-shot Learning(ZSL) attains knowledge transfer from seen classes to unseen classes by exploring auxiliary category information, which is a promising yet difficult research topic. In this field, Audio-Visual Generalized Zero-Shot Learning~(AV-GZSL) has aroused researchers' great interest in which intricate relations within triple modalities~(audio, video,

  10. Jiyu Chen, Shuang Peng, Daxiong Luo, Fan Yang

    Transformer-based large language models (LLMs) encounter challenges in processing long sequences on edge devices due to the quadratic complexity of attention mechanisms and growing memory demands from Key-Value (KV) cache. Existing KV cache optimizations struggle with irreversible token eviction in long-output tasks, while alternative sequence modeling archi

  11. Belle Collaboration, K. Uno, K. Hayasaka, K. Inami

    We report a search for the lepton-flavor-violating decays $\tau^{\pm}\to\ell^{\pm}\alpha$~($\ell=e,\mu$), where $\alpha$ is an undetected spin-0 particle, such as an axion-like particle using $736\times10^{6}$ tau lepton pairs collected by the Belle detector at the KEKB asymmetric-energy $e^{+}e^{-}$ collider. We find no evidence of signal and obtain the mos

  12. Yunhong Min, Daehyeon Choi, Kyeongmin Yeo, Jihyun Lee

    We introduce ORIGEN, the first zero-shot method for 3D orientation grounding in text-to-image generation across multiple objects and diverse categories. While previous work on spatial grounding in image generation has mainly focused on 2D positioning, it lacks control over 3D orientation. To address this, we propose a reward-guided sampling approach using a

  13. Yang Liu, Feixiang Liu, Jiale Du, Xinbo Gao

    Convolutional neural networks and supervised learning have achieved remarkable success in various fields but are limited by the need for large annotated datasets. Few-shot learning (FSL) addresses this limitation by enabling models to generalize from only a few labeled examples. Transductive few-shot learning (TFSL) enhances FSL by leveraging both labeled an

  14. Anindya Sarkar, G. Vadivu

    This research proposes a cutting-edge ensemble deep learning framework for stock price prediction by combining three advanced neural network architectures: The particular areas of interest for the research include but are not limited to: Variational Autoencoder (VAE), Transformer, and Long Short-Term Memory (LSTM) networks. The presented framework is aimed t

  15. Gargi Bakshi, Sujoy Bhore, Suraj Shetiya

    The outbreak of a pandemic, such as COVID-19, causes major health crises worldwide. Typical measures to contain the rapid spread usually include effective vaccination and strict interventions (Nature Human Behaviour, 2021). Motivated by such circumstances, we study the problem of limiting the spread of a disease over a social network system. In their seminal

  16. Yanting Peng, Zunyi Deng, Siyu Song, Gang Tang

    The recently reported auxetic piezoelectric effect, which acts as the electrical counterpart of the negative Poisson's ratio, is of significant technical importance for applications in acoustic wave devices. However, this electric auxetic effect has not yet been reported in perovskite systems. In this work, we employ first-principles calculations to investig

  17. Alexander Pushnitski, Sergei Treil

    A Hankel operator $\Gamma$ in $L^2(\mathbb{R}_+)$ is an integral operator with the integral kernel of the form $h(t+s)$, where $h$ is known as the kernel function. It is known that $\Gamma$ is positive semi-definite if and only if $h$ is the Laplace transform of a positive measure $\mu$ on $\mathbb{R}_+$. Thus, positive semi-definite Hankel operators $\Gamma

  18. Andrew Fowlie

    In collider physics, experiments are often based on counting the numbers of events in bins of a histogram. We present a new way to build and analyze statistical models that describe these experiments, based on the probabilistic programming language Stan and the HistFactory specification. A command-line tool transpiles HistFactory models into Stan code and da

  19. Hua-Wei Zhao, Yong Xie, Xinyao Huang, Guo-Feng Zhang

    Quantum batteries (QBs), harnessing quantum systems to transfer and store energy, have garnered substantial attention recently, enabling potentials in enhanced charging capacity, increased charging power, and device miniaturization. However, constrained by the weak interaction between the quantum nodes, the implementations of QB networks exhibit limited char

  20. Weicai Li, Tiejun Lv, Wei Ni, Jingbo Zhao

    Decentralized federated learning (D-FL) allows clients to aggregate learning models locally, offering flexibility and scalability. Existing D-FL methods use gossip protocols, which are inefficient when not all nodes in the network are D-FL clients. This paper puts forth a new D-FL strategy, termed Route-and-Aggregate (R&A) D-FL, where participating clients e

  21. Shadi Safaei Jazi, Ihar Faniayeu, Rafael Cichelero, Nikolai Kuznetsov

    The nonreciprocal magnetoelectric effect in Tellegen materials enables exotic phenomena such as axion-modified electrodynamics and fosters the development of magnet-free nonreciprocal media. As the nonreciprocal counterpart to the well-known chiral electromagnetic response, it offers a parallel framework in which many concepts developed for chiral materials

  22. Eleonora Di Nezza, Simon Jubert, Abdellah Lahdili

    In this paper we investigate the existence of metrics with weighted constant scalar curvature (wcscK for short) on a compact K\"ahler manifold $X$: this notion include constant scalar curvature K\"ahler metrics, weighted solitons, Calabi's extremal K\"ahler metrics and extremal metric on semisimple principal fibrations. We prove that the coercivity of the we

  23. Jianghao Lin, Peng Du, Jiaqi Liu, Weite Li

    E-commerce has revolutionized retail, yet its traditional workflows remain inefficient, with significant resource costs tied to product design and inventory. This paper introduces a novel system deployed at Alibaba that uses AI-generated items (AIGI) to address these challenges with personalized text-to-image generation for e-commerce product design. AIGI en

  24. Rahul Raja, Arpita Vats

    This paper presents the winning submission of the RaaVa team to the AmericasNLP 2025 Shared Task 3 on Automatic Evaluation Metrics for Machine Translation (MT) into Indigenous Languages of America, where our system ranked first overall based on average Pearson correlation with the human annotations. We introduce Feature-Union Scorer (FUSE) for Evaluation, FU

  25. Kanako Esaki, Tadayuki Matsumura, Yang Shao, Hiroyuki Mizuno

    This paper proposes the e-person architecture for constructing a unified and incremental development of AI ethics. The e-person architecture takes the reduction of uncertainty through collaborative cognition and action with others as a unified basis for ethics. By classifying and defining uncertainty along two axes - (1) first, second, and third person persp

  26. Juwei Guan, Xiaolin Fang, Donghyun Kim, Haotian Gong

    Low-quality data often suffer from insufficient image details, introducing an extra implicit aspect of camouflage that complicates camouflaged object detection (COD). Existing COD methods focus primarily on high-quality data, overlooking the challenges posed by low-quality data, which leads to significant performance degradation. Therefore, we propose KRNet,

  27. Dailan He, Xiahong Wang, Shulun Wang, Guanglu Song

    Face swapping aims to seamlessly transfer a source facial identity onto a target while preserving target attributes such as pose and expression. Diffusion models, known for their superior generative capabilities, have recently shown promise in advancing face-swapping quality. This paper addresses two key challenges in diffusion-based face swapping: the prior

  28. Siddhartha Siddhiprada Bhoi, Arathi Arakala, Amy Beth Corman, Asha Rao

    Homomorphic Encryption (HE) allows secure and privacy-protected computation on encrypted data without the need to decrypt it. Since Shor's algorithm rendered prime factorisation and discrete logarithm-based ciphers insecure with quantum computations, researchers have been working on building post-quantum homomorphic encryption (PQHE) algorithms. Most of the

  29. Chanhyuk Lee, Jiho Choi, Chanryeol Lee, Donggyun Kim

    Model merging has emerged as a promising approach for unifying independently fine-tuned models into an integrated framework, significantly enhancing computational efficiency in multi-task learning. Recently, several SVD-based techniques have been introduced to exploit low-rank structures for enhanced merging, but their reliance on such manually designed rank

  30. Shuai Zhang, Jinliang Wang, Sujith Konandetails, Xu Wang

    Accurate and reliable selection of the appropriate acetabular cup size is crucial for restoring joint biomechanics in total hip arthroplasty (THA). This paper proposes a novel framework that integrates square-root velocity function (SRVF)-based elastic shape registration technique with an embedded deformation (ED) graph approach to reconstruct the 3D articul

  31. Bargava Subramanian, Naveen Kumarasami, Praveen Shastry, Kalyan Sivasailam

    Introduction: Bone health disorders like osteoarthritis and osteoporosis pose major global health challenges, often leading to delayed diagnoses due to limited diagnostic tools. This study presents an AI-powered system that analyzes knee X-rays to detect key pathologies, including joint space narrowing, sclerosis, osteophytes, tibial spikes, alignment issues

  32. Ruiqi Liu, Boyu Diao, Libo Huang, Hangda Liu

    Continual learning (CL) aims to learn new tasks while retaining past knowledge, addressing the challenge of forgetting during task adaptation. Rehearsal-based methods, which replay previous samples, effectively mitigate forgetting. However, research on enhancing the efficiency of these methods, especially in resource-constrained environments, remains limited

  33. Wanlong Wu, Xionghong He, Yanyu Ren, Diyu Shen

    The Cooling-Storage-Ring External-target Experiment (CEE) at Heavy Ion Research Facility in Lanzhou (HIRFL) is designed to study the properties of nuclear matter created in heavy-ion collisions at a few hundred MeV/$u$ to 1 GeV/$u$ beam energies, facilitating the research of quantum chromodynamics phase structure in the high-baryon-density region. Collective

  34. Jialun Pei, Zhangjun Zhou, Diandian Guo, Zhixi Li

    Intraoperative bleeding in laparoscopic surgery causes rapid obscuration of the operative field to hinder the surgical process and increases the risk of postoperative complications. Intelligent detection of bleeding areas can quantify the blood loss to assist decision-making, while locating bleeding points helps surgeons quickly identify the source of bleedi

  35. Zhenzhou Guo, Xiaotian Wang, Wenhong Wang, Gang Zhang

    Spin-polarized antiferromagnets (AFMs), including altermagnets, noncollinear AFMs, and two-dimensional layer-polarized AFMs, have emerged as transformative materials for next-generation spintronic and optoelectronic technologies. These systems uniquely combine spin-polarized electronic states with vanishing net magnetization, enabling ultrafast spin dynamics

  36. Minho Park, Sunghyun Park, Jungsoo Lee, Hyojin Park

    This paper addresses the challenge of data scarcity in semantic segmentation by generating datasets through text-to-image (T2I) generation models, reducing image acquisition and labeling costs. Segmentation dataset generation faces two key challenges: 1) aligning generated samples with the target domain and 2) producing informative samples beyond the trainin

  37. Min Cao, Yuxin Lu, Ziyin Zeng, Dong Yi

    Data plays a pivotal role in Text-Based Person Retrieval (TBPR) research. Mainstream research paradigm necessitates real-world person images with manual textual annotations for training models, posing privacy concerns and annotation burdens. Several pioneering efforts explore synthetic data generation, and yet still depend on real data as a foundation, inher

  38. Di Wang, Wei-Chen Fu

    Motivated by the first observation of CP violation in baryon decays, we study the topological amplitudes of bottom baryon decays in the $SU(3)_F$ limit. The topological diagrams of the charmless two-body decays of bottom baryons are presented in detail. The linear relations between topologies and $SU(3)$ irreducible amplitudes are derived through tensor cont

  39. Jian Wang, Yefan Wang

    We report the first analytic calculation of the Higgs boson decay width to gluons up to next-to-leading order in quantum chromodynamics including full dependence on the bottom and charm quark masses. The interference between top- and bottom/charm-quark-induced amplitudes provides unexpectedly large contributions because the large logarithms of the bottom (ch

  40. Woojung Han, Yeonkyung Lee, Chanyoung Kim, Kwanghyun Park

    Diffusion-based text-to-image (T2I) models have recently excelled in high-quality image generation, particularly in a training-free manner, enabling cost-effective adaptability and generalization across diverse tasks. However, while the existing methods have been continuously focusing on several challenges, such as "missing objects" and "mismatched attribute

  41. Tian Tian, Chen Qing, Yuxuan Liao, Jiajun Zhu

    In regular magneto-optical trap (MOT) systems, the delivery of six circularly polarized (CP) cooling beams requires complex and bulky optical arrangements including waveplates, mirrors, retroreflectors, etc. To address such technique challenges, we have proposed a beam delivery system for miniaturized MOT entirely based on meta-devices. The key component is

  42. Song Wang, Junhong Lin, Xiaojie Guo, Julian Shun

    While large language models (LLMs) have made significant progress in processing and reasoning over knowledge graphs, current methods suffer from a high non-retrieval rate. This limitation reduces the accuracy of answering questions based on these graphs. Our analysis reveals that the combination of greedy search and forward reasoning is a major contributor t

  43. Zhanke Zhou, Zhaocheng Zhu, Xuan Li, Mikhail Galkin

    Numerous applications of large language models (LLMs) rely on their ability to perform step-by-step reasoning. However, the reasoning behavior of LLMs remains poorly understood, posing challenges to research, development, and safety. To address this gap, we introduce landscape of thoughts (LoT), the first landscape visualization tool to inspect the reasoning

  44. Bowen Gao, Yanwen Huang, Yiqiao Liu, Wenxuan Xie

    The discovery of novel small molecule drugs remains a critical scientific challenge with far-reaching implications for treating diseases and advancing human health. Traditional drug development--especially for small molecule therapeutics--is a highly complex, resource-intensive, and time-consuming process that requires multidisciplinary collaboration. Recent

  45. Seong-Hyeon Hwang, Minsu Kim, Steven Euijong Whang

    We study model confidence calibration in class-incremental learning, where models learn from sequential tasks with different class sets. While existing works primarily focus on accuracy, maintaining calibrated confidence has been largely overlooked. Unfortunately, most post-hoc calibration techniques are not designed to work with the limited memories of old-

  46. Ning Liu, Sen Shen, Xiangrui Kong, Hongtao Zhang

    Multi-Agent Pathfinding is used in areas including multi-robot formations, warehouse logistics, and intelligent vehicles. However, many environments are incomplete or frequently change, making it difficult for standard centralized planning or pure reinforcement learning to maintain both global solution quality and local flexibility. This paper introduces a h

  47. Dinil Mon Divakaran

    Network traffic analysis using AI (machine learning and deep learning) models made significant progress over the past decades. Traffic analysis addresses various challenging problems in network security, ranging from detection of anomalies and attacks to countering of Internet censorship. AI models are also developed to expose user privacy risks as demonstra

  48. Yu-fan Zheng, Bin Chen

    In this work, we investigate possible supersymmetric extensions of the Carrollian algebra and the Carrollian conformal algebra in both $d=4$ and $d=3$. For the super-Carrollian algebra in $d=4$, we identify multiple admissible structures, depending on the representations of the supercharges with respect to the Carrollian rotation. Some of these structures ca

  49. Abdul Jabbar, Ethan Grooby, Jack Crozier, Alexander Gallon

    Congenital heart disease (CHD) is a critical condition that demands early detection, particularly in infancy and childhood. This study presents a deep learning model designed to detect CHD using phonocardiogram (PCG) signals, with a focus on its application in global health. We evaluated our model on several datasets, including the primary dataset from Bangl

  50. Hao Feng, Hao Sun, Wei Xie, Zhi Zuo

    While dynamic novel view synthesis from 2D videos has seen progress, achieving efficient reconstruction and rendering of dynamic scenes remains a challenging task. In this paper, we introduce Disentangled 4D Gaussian Splatting (Disentangled4DGS), a novel representation and rendering pipeline that achieves real-time performance without compromising visual fid

  51. Mingyuan Hong

    Quantum plasmonics explores how light interacts with collective charge oscillations at metal-dielectric interfaces, enabling strong confinement and enhanced quantum effects at the nanoscale. While traditional quantum optics focuses on single photons, this thesis explores an intermediate regime - multiparticle quantum optics - where classical light, analyzed

  52. Mingxiang Li, Josep M. Jornet, Daniel M. Mittleman, Chong Han

    The terahertz frequency band, ranging from 0.1 to 10 THz, offers extensive spectral resources for next-generation wireless communication systems. To compensate for the limited transmission power of terahertz transceivers and the significant propagation losses in terahertz channels, high-gain directional antennas are essential. Dynamic beam manipulation is th

  53. Chao Song, Kai Wang, Yuanyuan Zhang, Guodong Zhou

    This paper is the second in a series dedicated to the operadic study of Nijenhuis structures, focusing on Nijenhuis Lie algebras and Nijenhuis geometry. We introduce the concept of homotopy Nijenhuis Lie algebras and establish that the differential graded (=dg) operad $\mathfrak{NjL}_{\infty}$ governing these structures serves as the minimal model of the ope

  54. Zekai Liu, Xiaoqi Li

    Cryptocurrency is a novel exploration of a form of currency that proposes a decentralized electronic payment scheme based on blockchain technology and cryptographic theory. While cryptocurrency has the security characteristics of being distributed and tamper-proof, increasing market demand has led to a rise in malicious transactions and attacks, thereby expo

  55. Yingjie Fan, Xuewen Liu, Ning Zhou

    The neutrino floor, a theoretical sensitivity limit for dark matter (DM) direct detections, is being redefined as the boundary of a dynamic ``neutrino fog", where neutrino signals become inevitable, obscuring DM detection due to the statistical and systematic uncertainties. This study provides the first site-specific analysis of the neutrino floor at China J

  56. Jae-Young Yim, Dongwook Kim, Jae-Young Sim

    Large-scale datasets are usually required to train deep neural networks, but it increases the computational complexity hindering the practical applications. Recently, dataset distillation for images and texts has been attracting a lot of attention, that reduces the original dataset to a synthetic dataset to alleviate the computational burden of training whil

  57. Jackson A. Mickley, Waseem Kamleh, Derek B. Leinweber

    The importance of examining the structure of centre-vortex matter in the ground-state fields of nonabelian gauge-field theory has been demonstrated in the recent centre-vortex based discovery of a second finite-temperature transition in QCD associated with quark deconfinement. This signals the presence of a new phase of ground-state field structure between t

  58. Yuxuan Li, Vijay Veerabadran, Michael L. Iuzzolino, Brett D. Roads

    We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D dataset to benchmark the ability to predict a camera wearer's goals, beliefs, and next actions. We study the performance of both humans and state

  59. Joshua Krook

    AI risks are typically framed around physical threats to humanity, a loss of control or an accidental error causing humanity's extinction. However, I argue in line with the gradual disempowerment thesis, that there is an underappreciated risk in the slow and irrevocable decline of human autonomy. As AI starts to outcompete humans in various areas of life, a

  60. Prashant Mahajan

    As Artificial Intelligence (AI) systems increasingly permeate caregiving, educational, and emotionally sensitive domains, there is a growing need to assess national readiness beyond infrastructure and innovation capacity. Existing indices such as the Stanford AI Index (2024), overlooked relational, ethical, and cultural dimensions essential to human centered

  61. Saleh Sakib Ahmed, Rashed Uz Zzaman, Saifur Rahman Jony, Faizur Rahman Himel

    Long-term groundwater level (GWL) measurement is vital for effective policymaking and recharge estimation using annual maxima and minima. However, current methods prioritize short-term predictions and lack multi-year applicability, limiting their utility. Moreover, sparse in-situ measurements lead to reliance on low-resolution satellite data like GLDAS as th

  62. Matthew Babbitt

    Summability has been a central object of study in difference algebra over the past half-century. It serves as a cornerstone of algebraic methods to study linear recurrences over various fields of coefficients and with respect to various kinds of difference operators. Recently, Dreyfus, Hardouin, Roques, and Singer introduced a notion of elliptic orbital resi

  63. Rong Du, Yuhang Zhou

    There is a long-standing conjecture which states that every uniform algebraic vector bundle of rank $r<2n$ on the $n$-dimensional projective space $\mathbb{P}^n$ over an algebraically closed field of characteristic $0$ is homogeneous. This conjecture is valid for $n\leq3$. In this paper, we classify all uniform vector bundles of rank $r<8$ over $\mathbb{P}^4

  64. Mitchell T. Dennis, Esther M. Hu, Lennox L. Cowie

    We present the result of two binary classifier ensembled neural networks to identify catastrophic outliers for photo-z estimates within the COSMOS field utilizing only 8 and 5 photometric band passes, respectively. Our neural networks can correctly classify 55.6% and 33.3% of the true positives with few to no false positives. These methods can be used to red

  65. Guansen Wang, Bing-Yu Su, Lei Zu, Lei Feng

    If sub-GeV Dark matter(DM) annihilates to the charged particles such as $e^+ e^-$, $\mu^+ \mu^-$, or $\pi^+ \pi^-$, it generates an additional source of electrons and positrons in the cosmic ray (CR) population within our Milky Way. During propagation, these secondary electrons and positrons undergo reacceleration processes, boosting their energies to the Ge

  66. Sohail Reddy

    Characterizing non-Markovian quantum dynamics is essential for accurately modeling open quantum systems, particularly in near-term quantum technologies. In this work, we develop a structure-preserving approach to characterizing non-Markovian evolution using the time-convolutionless (TCL) master equation, considering both linear and nonlinear formulations. To

  67. Othmane Benhaida, El Hassan Saidi, L. B. Drissi

    Haldane's tight-binding model, which describes a Chern insulator in a two-dimensional hexagonal lattice, exhibits quantum Hall conductivity without an external magnetic field. Here, we explore an $\alpha -T_{3}$ lattice subjected to circularly polarized off-resonance light. This lattice, composed of two sublattices (A and B) and a central site (C) per unit c

  68. Tim Rolff, Jurik Karimian, Niklas Hypki, Susanne Schmidt

    A considerable part of the performance of today's large language models (LLM's) and multimodal large language models (MLLM's) depends on their tokenization strategies. While tokenizers are extensively researched for textual and visual input, there is no research on tokenization strategies for gaze data due to its nature. However, a corresponding tokenization

  69. Papa Abdou Karim Karou Diallo, Amal Zouaq

    Translating natural language questions into SPARQL queries enables Knowledge Base querying for factual and up-to-date responses. However, existing datasets for this task are predominantly template-based, leading models to learn superficial mappings between question and query templates rather than developing true generalization capabilities. As a result, mode

  70. Sungyu Jeong, Won Joon Choi, Junung Choi, Anik Biswas

    We propose a UNet-based foundation model and its self-supervised learning method to address two key challenges: 1) lack of qualified annotated analog layout data, and 2) excessive variety in analog layout design tasks. For self-supervised learning, we propose random patch sampling and random masking techniques automatically to obtain enough training data fro

  71. Qingtang Su, Siwei Wang

    We study the two-dimensional gravity water waves with a one-dimensional interface with small initial data. Our main contributions include the development of two novel localization lemmas and a Transition-of-Derivatives method, which enable us to reformulate the water wave system into the following simplified structure: $$(D_t^2-iA\partial_{\alpha})\theta=i\f

  72. Yifan Zhang, Dave Towey, Matthew Pike, Quang-Hung Luu

    Context: This paper provides an in-depth examination of the generation and evaluation of Metamorphic Relations (MRs) using GPT models developed by OpenAI, with a particular focus on the capabilities of GPT-4 in software testing environments. Objective: The aim is to examine the quality of MRs produced by GPT-3.5 and GPT-4 for a specific System Under Test (SU

  73. Chang Cai, Xiaojun Yuan, Ying-Jun Angela Zhang

    Message passing algorithms have been tailored for compressive imaging applications by plugging in different types of off-the-shelf image denoisers. These off-the-shelf denoisers mostly rely on some generic or hand-crafted priors for denoising. Due to their insufficient accuracy in capturing the true image prior, these methods often fail to produce satisfacto

  74. Jiacheng Xie, Hua-Chieh Shao, You Zhang

    Time-resolved CBCT imaging, which reconstructs a dynamic sequence of CBCTs reflecting intra-scan motion (one CBCT per x-ray projection without phase sorting or binning), is highly desired for regular and irregular motion characterization, patient setup, and motion-adapted radiotherapy. Representing patient anatomy and associated motion fields as 3D Gaussians

  75. Changchang Sun, Gaowen Liu, Charles Fleming, Yan Yan

    Conditional diffusion models have gained increasing attention since their impressive results for cross-modal synthesis, where the strong alignment between conditioning input and generated output can be achieved by training a time-conditioned U-Net augmented with cross-attention mechanism. In this paper, we focus on the problem of generating music synchronize

  76. Syrine Belakaria, Joshua Kazdan, Charles Marx, Chris Cundy

    Reinforcement learning from human feedback (RLHF) has become a cornerstone of the training and alignment pipeline for large language models (LLMs). Recent advances, such as direct preference optimization (DPO), have simplified the preference learning step. However, collecting preference data remains a challenging and costly process, often requiring expert an

  77. Hongmei Yin, Tingliang Feng, Fan Lyu, Fanhua Shang

    In this work, we focus on continual semantic segmentation (CSS), where segmentation networks are required to continuously learn new classes without erasing knowledge of previously learned ones. Although storing images of old classes and directly incorporating them into the training of new models has proven effective in mitigating catastrophic forgetting in c

  78. Zhipeng Lu

    We focus on establishing the foundational paradigm of a novel optimization theory based on convolution with convex kernels. Our goal is to devise a morally deterministic model of locating the global optima of an arbitrary function, which is distinguished from most commonly used statistical models. Limited preliminary numerical results are provided to test th

  79. Costain Nachuma, Md Mosharaf Hossan, Asif Kamal Turzo, Minhaz F. Zibran

    This study investigates vulnerabilities within the Maven ecosystem by analyzing a comprehensive dataset of 14,459,139 releases. Our analysis reveals the most critical weaknesses that pose significant threats to developers and their projects as they look to streamline their development tasks through code reuse. We show risky weaknesses, those unique to Maven,

  80. Jason E. McDermott, William C. Nelson, Amy E. Zimmerman, Winston Anthony

    The introduction of non-native organisms into complex microbiome communities holds enormous potential to benefit society. However, microbiome engineering faces several challenges including successful establishment of the organism into the community, its persistence in the microbiome to serve a specified purpose, and constraint of the organism and its activit

  81. Toma Masaki, Kanta Tachibana

    This study proposes a novel framework for long-term electricity demand prediction based solely on historical consumption data, without relying on external variables such as temperature or economic indicators. The method combines Non-negative Tensor Factorization (NTF) to extract low-dimensional temporal features from multi-way electricity usage data, with a

  82. Dayou Luo, Yue Yu, Maryam Fazel, Behçet Açıkmeşe

    We propose Newton-PIPG, an efficient method for solving quadratic programming (QP) problems arising in optimal control, subject to additional set constraints. Newton-PIPG integrates the Proportional-Integral Projected Gradient (PIPG) method with the Newton method, thereby achieving both global convergence and local quadratic convergence. The PIPG method, an

  83. Meng Zhang, Bin Yue, Yidong Xu, Andrea Ferrara

    JWST reveals numerous high-z galaxies and Supermassive Black Holes (SMBHs), suggesting that stars and SMBH seeds formation at $z \gtrsim 10$ may be more efficient than previously derived. One popular SMBH seed scenario is the Direct Collapse Black Holes (DCBHs) formed in pristine atomic-cooling halos irradiated by nearby galaxies. Therefore, the efficient st

  84. Josh Schipper, Radnya Mukhedkar, Neville Watson, Veerabrahmam Bathini

    This work identifies the general approach for linearising any power system component in the harmonic domain, that is with respect to its Fourier series coefficients. This ability enables detailed harmonic analysis, and is key as more power electronic devices inject harmonic currents into the power system to its shared detriment. The general approach requires

  85. Yuki Uematsu

    Microbubble solutions have a wide range of industrial applications, including heat transfer, agriculture, and water treatment. Therefore, understanding and controlling the size variation of bubbles is critical. In this study, we develop a theoretical framework for Ostwald ripening in buoyancy-driven microbubbles by introducing a height-dependent size distrib

  86. Qiang Xiong, Tanda Li, Jie Yu, Zhiwen Li

    Long-period variables (LPVs) are high-luminosity red giants or supergiants with pulsation periods ranging from days to years. Many LPVs in the Large Magellanic Cloud (LMC) and Galactic Bulge (BLG) have been continuously observed over a time span of 26 years by the Optical Gravitational Lensing Experiment (OGLE) survey. Using OGLE-IV data, we applied Gaussian

  87. BESIII Collaboration, M. Ablikim, M. N. Achasov, P. Adlarson

    The strong-phase differences between $D^0\to K_{S/L}^0\pi^+\pi^-$ and $\bar{D}^0\to K_{S/L}^0\pi^+\pi^-$ decays are one of the most important inputs in measuring the $C\!P$ violating angle $\gamma$ via $B^- \to D K^-$ decays. They also play a key role in studies of charm mixing and indirect $C\!P$ violation. In this paper, the strong-phase differences are de

  88. Ivan Beleacov

    Automated construction is one of the most promising areas that can improve efficiency, reduce costs and minimize errors in the process of building construction. In this paper, a comparative analysis of three neural network models for semantic segmentation, U-Net(light), LinkNet and PSPNet, is performed. Two specialized datasets with images of houses built fr

  89. Amr Alshatnawi, Remi Sampaleanu, David Liebovitz

    Artificial Intelligence (AI) has been advancing rapidly and with the advent of large language models (LLMs) in late 2022, numerous opportunities have emerged for adopting this technology across various domains, including medicine. These innovations hold immense potential to revolutionize and modernize medical education. Our research project leverages large l

  90. Peng Lin, Haopeng Yang, Gui Gui, Mengxiang Zeng

    In this paper, the scheduling problems of landing and takeoff aircraft on a same runway and on dual runways are addressed. In contrast to the approaches based on mixed-integer optimization models in existing works, our approach focuses on the minimum separation times between aircraft by introducing some necessary assumptions and new concepts including releva

  91. Hunter D. Ellis, Botong Li, Haoyu Xie, Jichao Fan

    $\beta$-Ga$_2$O$_3$ is gaining attention as a promising semiconductor for next-generation high-power, high-efficiency, and high-temperature electronic devices, thanks to its exceptional material properties. However, challenges such as the lack of viable p-type doping have hindered its full potential, particularly in the development of ambipolar devices. This

  92. Puzhen Yuan, Angyuan Ma, Yunchao Yao, Huaxiu Yao

    Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition. Existing methods typically rely on prior environmental knowledge or carefully designed task-specific prompts, making them struggle with dynamic scene

  93. Tharun Anand, Siva Sankar Sajeev, Pravin Nair

    With rapid advancements in generative modeling, deepfake techniques are increasingly narrowing the gap between real and synthetic videos, raising serious privacy and security concerns. Beyond traditional face swapping and reenactment, an emerging trend in recent state-of-the-art deepfake generation methods involves localized edits such as subtle manipulation

  94. Protyay Dey, Rejoy Chakraborty, Abhilasha S. Jadhav, Kapil Rana

    Source camera model identification (SCMI) plays a pivotal role in image forensics with applications including authenticity verification and copyright protection. For identifying the camera model used to capture a given image, we propose SPAIR-Swin, a novel model combining a modified spatial attention mechanism and inverted residual block (SPAIR) with a Swin

  95. Chenya Huang, Zhidong Li, Fang Chen, Bin Liang

    Real estate appraisal has undergone a significant transition from manual to automated valuation and is entering a new phase of evolution. Leveraging comprehensive attention to various data sources, a novel approach to automated valuation, multimodal machine learning, has taken shape. This approach integrates multimodal data to deeply explore the diverse fact

  96. Muhammad Usama, Haris N. Koutsopoulos, Zhengbing He, Lijiao Wang

    Driving cycles are a set of driving conditions and are crucial for the existing emission estimation model to evaluate vehicle performance, fuel efficiency, and emissions, by matching them with average speed to calculate the operating modes, such as braking, idling, and cruising. While existing emission estimation models, such as the Motor Vehicle Emission Si

  97. John Mellnik, Jack Scannell

    Consider two similar drug companies with access to similar chemical libraries and synthesis methods, who each run an R&D program. The programs have the same number of stages, which each take the same amount of time, with the same costs, with the same historic stepwise progression rates, and which aim to address the same therapeutic indication. Now let us sup

  98. Alice Qian Zhang, Jina Suh, Mary L. Gray, Hong Shen

    As artificial intelligence (AI) systems become increasingly embedded in critical societal functions, the need for robust red teaming methodologies continues to grow. In this forum piece, we examine emerging approaches to automating AI red teaming, with a particular focus on how the application of automated methods affects human-driven efforts. We discuss the

  99. Yazhou Zhang, Qimeng Liu, Qiuchi Li, Peng Zhang

    Evaluating the value alignment of large language models (LLMs) has traditionally relied on single-sentence adversarial prompts, which directly probe models with ethically sensitive or controversial questions. However, with the rapid advancements in AI safety techniques, models have become increasingly adept at circumventing these straightforward tests, limit

  100. Atsuki Kumashita, Hiroo Tajiri, Jun Usami, Yu Yamane

    We observed surface X-ray diffraction from He-4 submonolayers adsorbed on a single-surface graphite using synchrotron X-rays. Time evolutions of scattering intensities along the crystal truncation rod (CTR) were observed even after reaching the base low temperature in a selected condition of sample preparation. Our simulations for CTR scatterings based on th