Skip to content

December 2024 arXiv papers — page 38

Showing 3,7013,800 of 20,868 papers

  1. Yaoyun Zhang, Xuenan Xu, Mengyue Wu

    The video-to-audio (V2A) generation task has drawn attention in the field of multimedia due to the practicality in producing Foley sound. Semantic and temporal conditions are fed to the generation model to indicate sound events and temporal occurrence. Recent studies on synthesizing immersive and synchronized audio are faced with challenges on videos with mo

  2. Cong Li, Qingqing Long, Yuanchun Zhou, Meng Xiao

    Large language models (LLMs) have demonstrated remarkable advancements, primarily due to their capabilities in modeling the hidden relationships within text sequences. This innovation presents a unique opportunity in the field of life sciences, where vast collections of single-cell omics data from multiple species provide a foundation for training foundation

  3. Taisei Yamanaka, Yoshihiko Ihara, Satoru Hayami

    We theoretically propose the emergence of nonlinear nonreciprocal conductivity in centrosymmetric paramagnetic systems when a spatially gradient magnetic field is externally applied. The key essence lies in the appearance of magnetic toroidal dipole moment under the gradient field that breaks both spatial inversion and time-reversal symmetries. By analyzing

  4. Zhijian Chen, Chuan Hu, Min Wu, Qingqing Long

    Emerging topics in biomedical research are continuously expanding, providing a wealth of information about genes and their function. This rapid proliferation of knowledge presents unprecedented opportunities for scientific discovery and formidable challenges for researchers striving to keep abreast of the latest advancements. One significant challenge is nav

  5. Zhiheng Liu, Ka Leong Cheng, Qiuyu Wang, Shuzhe Wang

    Missing values remain a common challenge for depth data across its wide range of applications, stemming from various causes like incomplete data acquisition and perspective alteration. This work bridges this gap with DepthLab, a foundation depth inpainting model powered by image diffusion priors. Our model features two notable strengths: (1) it demonstrates

  6. Rahul Gupta, Judith Racusin, Vladimir Lipunov, Y. -D. Hu

    Robotic telescope networks play an important role in capturing early and bright optical afterglows, providing critical insights into the energetics and emission mechanisms of GRBs. In this study, we analyze GRB 230204B, an exceptionally energetic and multi-pulsed long GRB, detected by the Fermi GBM and MAXI detectors, with an isotropic equivalent gamma-ray e

  7. Yusuke Ide, Joshua Tanner, Adam Nohejl, Jacob Hoffman

    Multiword expressions (MWEs) refer to idiomatic sequences of multiple words. MWE identification, i.e., detecting MWEs in text, can play a key role in downstream tasks such as machine translation, but existing datasets for the task are inconsistently annotated, limited to a single type of MWE, or limited in size. To enable reliable and comprehensive evaluatio

  8. Shuhao Han, Haotian Fan, Jiachen Fu, Liang Li

    Recently, Text-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated metrics have emerged to evaluate the image-text alignment capabilities of generative models. However, the performance comparison among these automated metrics is limited by existing small datasets. Additionally, these datasets lack the capa

  9. Xiao Guo, Manh Tran, Jiaxin Cheng, Xiaoming Liu

    The text-to-image (T2I) personalization diffusion model can generate images of the novel concept based on the user input text caption. However, existing T2I personalized methods either require test-time fine-tuning or fail to generate images that align well with the given text caption. In this work, we propose a new T2I personalization diffusion model, Dense

  10. Zhen Sun, Zongmin Zhang, Xinyue Shen, Ziyi Zhang

    Social media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs). However, the misuse of AIGTs could have profound implications for public opinion, such as spreading misinformation and manipulating narratives. Despite its importance, it remains unclear how prevalent AIGTs are on social media. To address this gap, this paper aims to qu

  11. Robinson Umeike, Thang Dao, Shane Crawford

    Post-disaster assessments of buildings and infrastructure are crucial for both immediate recovery efforts and long-term resilience planning. This research introduces an innovative approach to automating post-disaster assessments through advanced deep learning models. Our proposed system employs state-of-the-art computer vision techniques (YOLOv11 and ResNet5

  12. Jiancheng Wu, Qingwen Wu, Kaixing Lu, Xinwu Cao

    The geometry and kinematics of the broad-line region (BLR) in AGNs are still unclear, which is crucial for studying the physics and evolution of supermassive black holes (SMBHs) and AGNs. The broad-line profile provides valuable information on BLR geometry and kinematics. In this work, we explore the evolution of line profiles in variable AGNs based on the B

  13. Yingying Ma, Wei Lan, Chenlei Leng, Ting Li

    The social characteristics of players in a social network are closely associated with their network positions and relational importance. Identifying those influential players in a network is of great importance as it helps to understand how ties are formed, how information is propagated, and, in turn, can guide the dissemination of new information. Motivated

  14. Ruipu Li, Alexander Rodríguez

    We introduce a neural network conformal prediction method for time series that enhances adaptivity in non-stationary environments. Our approach acts as a neural controller designed to achieve desired target coverage, leveraging auxiliary multi-view data with neural network encoders in an end-to-end manner to further enhance adaptivity. Additionally, our mode

  15. Veronica Santos, Bruno Cuconato

    Graphs are the most suitable structures for modeling objects and interactions in applications where component inter-connectivity is a key feature. There has been increased interest in graphs to represent domains such as social networks, web site link structures, and biology. Graph stores recently rose to prominence along the NoSQL movement. In this work we w

  16. Youngmoon Jung, Jinyoung Lee, Seungjin Lee, Myunghun Jung

    Recent advances in flexible keyword spotting (KWS) with text enrollment allow users to personalize keywords without uttering them during enrollment. However, there is still room for improvement in target keyword performance. In this work, we propose a novel few-shot transfer learning method, called text-aware adapter (TA-adapter), designed to enhance a pre-t

  17. Wen Wen, Qiang Zhou, Yu Xi, Haoyu Li

    In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech enhancement system remains challenging, especially in extremely low signal-to-noise ratio (SNR) conditions. To tackle this issue, we propose a triple-steering spatial selection metho

  18. Rui Ai, Boxiang Lyu, Zhaoran Wang, Zhuoran Yang

    How much value does a dataset or a data production process have to an agent who wishes to use the data to assist decision-making? This is a fundamental question towards understanding the value of data as well as further pricing of data. This paper develops an approach for capturing the instrumental value of data production processes, which takes two key fact

  19. Chengpeng Fu, Xiaocheng Feng, Yichong Huang, Wenshuai Huo

    The in-image machine translation task involves translating text embedded within images, with the translated results presented in image format. While this task has numerous applications in various scenarios such as film poster translation and everyday scene image translation, existing methods frequently neglect the aspect of consistency throughout this proces

  20. Benjamin Laufer, Manish Raghavan, Solon Barocas

    Disparate impact doctrine offers an important legal apparatus for targeting discriminatory data-driven algorithmic decisions. A recent body of work has focused on conceptualizing one particular construct from this doctrine: the less discriminatory alternative, an alternative policy that reduces disparities while meeting the same business needs of a status qu

  21. Tiantian Mu, Jun-e Feng, Biao Wang

    This paper exploits bisimulation relations, generated by extracting the concept of morphisms between algebraic structures, to analyze set stabilization of Boolean control networks with lower complexity. First, for two kinds of bisimulation relations, called as weak bisimulation and strong bisimulation relations, a novel verification method is provided by con

  22. Le Dong, Qixuan Cao, Lei Pu, Fangfang Wu

    ERVD: An Efficient and Robust ViT-Based Distillation Framework for Remote Sensing Image Retrieval

  23. Binrui Zeng, Bin Ji, Xiaodong Liu, Jie Yu

    As Large Language Models (LLMs) demonstrate exceptional performance across various domains, deploying LLMs on edge devices has emerged as a new trend. Quantization techniques, which reduce the size and memory requirements of LLMs, are effective for deploying LLMs on resource-limited edge devices. However, existing one-size-fits-all quantization methods often

  24. Huanhuan Wei, Jing Tang, Yuangang Deng

    High-spin quantum systems, endowed with rich internal degrees of freedom, constitute a promising platform for manipulating high-quality $n$-photon states. In this study, we explore $n$-photon bundles emission by constructing a high-spin Jaynes-Cummings model (JCM) within a single-mode cavity interacting with a single spin-$3/2$ atom. Our analysis reveals tha

  25. Weiqi Zhou

    Let $p$ be a prime number, it is shown that tiling and spectral sets coincide in $\mathbb Z_{p^2}\times\mathbb Z_{p^2}$ by considering equivalently symplectic spectral pairs. The main approach is still to analyze the zero set of the Fourier transform. The zero set of the symplectic Fourier transform differs from the zero set of the usual Fourier transform by

  26. Yuru Wang, Pei Liu, Songtao Wang, Zehan Zhang

    Open-world 3D scene understanding is a critical challenge that involves recognizing and distinguishing diverse objects and categories from 3D data, such as point clouds, without relying on manual annotations. Traditional methods struggle with this open-world task, especially due to the limitations of constructing extensive point cloud-text pairs and handling

  27. Jianfei Xu, Rui Zhang, Junhui Fan

    The study takes the social media industry as its research subject and examines the impact of scientific innovation capabilities on profit distribution within the value chain of the social media industry. It proposes a specific solution to the profit distribution problem using an improved Shapley value method. Additionally, the AHP (Analytic Hierarchy Process

  28. Ziye Zheng, Jiajing Wu, Dan Lin, Quanzhong Li

    As the number of blockchain platforms continues to grow, the independence of these networks poses challenges for transferring assets and information across chains. Cross-chain bridge technology has emerged to address this issue, establishing communication protocols to facilitate cross-chain interaction of assets and information, thereby enhancing user experi

  29. Mingyue Guo, Zhenhua Shi

    A generalized Camassa-Holm equation, which describes pseudospherical surfaces, is considered. Using geometric methods, it is demonstrated that the equation is geometrically integrable. Additionally, an infinite hierarchy of conservation laws is derived. Furthermore, the paper delves into the investigation of the problem of locally isometric immersion into th

  30. Thomas Creutzig, Robert McRae, Florencia Orosz Hunziker, Jinwei Yang

    We show that the category of $C_1$-cofinite modules for the universal $N=1$ super Virasoro vertex operator superalgebra $\mathcal{S}(c,0)$ at any central charge $c$ is locally finite and admits the vertex algebraic braided tensor category structure of Huang-Lepowsky-Zhang. For central charges $c^{\mathfrak{ns}}(t)=\frac{15}{2}-3(t+t^{-1})$ with $t\notin\math

  31. Shiqi Yin, Min Dong

    The main challenges in designing downlink coordinated multicast beamforming in massive multiple-input multiple output (MIMO) cellular networks are the complex computational solutions and significant fronthaul overhead for centralized coordination. This paper proposes a coordinated multicast beamforming solution that is both computation and communication effi

  32. Qian Tao, Xiyuan Wang, Muhan Zhang, Shuxian Hu

    Graph neural networks (GNNs) have become a prevalent framework for graph tasks. Many recent studies have proposed the use of graph convolution methods over the numerous subgraphs of each graph, a concept known as subgraph graph neural networks (subgraph GNNs), to enhance GNNs' ability to distinguish non-isomorphic graphs. To maximize the expressiveness, subg

  33. Zhaohui Jin, Yi Shuai, Yongcheng Li, Lingcong Cai

    The early detection of glottic carcinoma is critical for improving patient outcomes, as it enables timely intervention, preserves vocal function, and significantly reduces the risk of tumor progression and metastasis. However, the similarity in morphology between glottic carcinoma and vocal cord dysplasia results in suboptimal detection accuracy. To address

  34. Yiming Wang, Jiahao Chen, Qingming Li, Tong Zhang

    As text-to-image (T2I) models advance and gain widespread adoption, their associated safety concerns are becoming increasingly critical. Malicious users exploit these models to generate Not-Safe-for-Work (NSFW) images using harmful or adversarial prompts, underscoring the need for effective safeguards to ensure the integrity and compliance of model outputs.

  35. Si Wang, Guoqiang Xiao

    Array structures based on the fourth-order difference co-array (FODCA) provide more degrees of freedom (DOF). However, since the growth of DOF is limited by a single case of fourth-order cumulant in FODCA, this paper aims to design a sparse linear array (SLA) with higher DOF via exploring different cases of fourth-order cumulants. This paper presents a mathe

  36. Xiaoyang Hu, Richard L. Lewis

    Cognitive tasks originally developed for humans are now increasingly used to study language models. While applying these tasks is often straightforward, interpreting their results can be challenging. In particular, when a model underperforms, it is often unclear whether this results from a limitation in the cognitive ability being tested or a failure to unde

  37. Hongyi He, Haoyue Tang, Jiayu Pan, Jintao Wang

    In this paper, we study a system in which a sensor forwards status updates to a receiver through an error-prone channel, while the receiver sends the transmission results back to the sensor via a reliable channel. Both channels are subject to random delays. To evaluate the timeliness of the status information at the receiver, we use the Age of Information (A

  38. Tomoatsu Edagawa, Kazuki Yoshida, Shoichiro Kawase, Kazuyuki Ogata

    It is shown that longitudinally polarized protons can be used to induce chirality in the final states of the $(\vec{p},pN)$ reaction at intermediate energies, when there exist three final-state particles with non-coplanar momentum vectors. The analyzing power $A_z$ is proposed as a measure of this effect. Theoretical descriptions to obtain $A_z$ based on an

  39. Amol Aggarwal, Ivan Corwin, Milind Hegde

    We consider the stochastic six-vertex (S6V) model and asymmetric simple exclusion process (ASEP) under general initial conditions which are bounded below lines of arbitrary slope at $\pm\infty$. We show under Kardar-Parisi-Zhang (KPZ) scaling of time, space, and fluctuations that the height functions of these models converge to the KPZ fixed point. Previousl

  40. Hao Wen, Shizuo Tian, Borislav Pavlov, Wenjie Du

    Large language models (LLMs) have brought exciting new advances to mobile UI agents, a long-standing research field that aims to complete arbitrary natural language tasks through mobile UI interactions. However, existing UI agents usually demand powerful large language models that are difficult to be deployed locally on end-users' devices, raising huge conce

  41. David P. Carcamo, Christopher W. Lynn

    As experiments advance to record from tens of thousands of neurons, statistical physics provides a framework for understanding how collective activity emerges from networks of fine-scale correlations. While modeling these populations is tractable in loop-free networks, neural circuitry inherently contains feedback loops of connectivity. Here, for a class of

  42. Nguyen Ngoc Hai, Le Dung Muu, Nguyen Van Quy

    We consider class of equilibrium models including the implicit Walras supply-demand and competitive models. Such a model in this class, in general, is ill-posed. We formulate such a model in the form a variational inequality having certain monotonicity property which allow us to describe a regularization algorithm avoiding the ill-posedness based upon the bi

  43. Esteban Andruchow, Eduardo Chiumiento

    Let H be a separable complex Hilbert space. Denote by Gr(H) the Grassmann manifold of H. We study the following sets of pairs of elements in Gr(H): Delta={(S,T) in Gr(H) x Gr(H): there exists Z in Gr(H) such that S\dot{+} Z=T \dot{+} Z=H }, which are pairs of subspaces that have a common complement, and Gamma={(S,T) in Gr(H) x Gr(H): (S,T) does not belong to

  44. Peifu Liu, Tingfa Xu, Guokai Shi, Jingxuan Xu

    Hyperspectral salient object detection (HSOD) aims to extract targets or regions with significantly different spectra from hyperspectral images. While existing deep learning-based methods can achieve good detection results, they generally necessitate pixel-level annotations, which are notably challenging to acquire for hyperspectral images. To address this i

  45. Mingming Zhang, Zhiqing Xiao, Guoshan Lu, Sai Wu

    Tabular data, which accounts for over 80% of enterprise data assets, is vital in various fields. With growing concerns about privacy protection and data-sharing restrictions, generating high-quality synthetic tabular data has become essential. Recent advancements show that large language models (LLMs) can effectively gener-ate realistic tabular data by lever

  46. Gui Ling, Ziyang Wang, Yuliang Yan, Qingwen Liu

    Large language models (LLMs) have garnered significant attention for their remarkable capabilities across various domains, whose vast parameter scales present challenges for practical deployment. Structured pruning is an effective method to balance model performance with efficiency, but performance restoration under computational resource constraints is a pr

  47. Ning Zhang

    We present necessary and sufficient conditions for a group homomorphism between spaces of smooth sections of Lie group bundles to be a weighted composition operator. These results provide new insights into a wide range of problems related to weighted composition operators. Specifically, we prove that the algebraic structure of the space of smooth sections of

  48. Akshay Sathiya, Rohit Pandey

    Today, several people and organizations rely on cloud platforms. The reliability of cloud platforms depends heavily on the performance of their internal programs (agents). To better prevent regressions in cloud platforms, the design of pre-production testing environments (that test new agents, new hardwares, and other changes) must take into account the dive

  49. Jaafar Gaber

    This paper introduces a novel cryptographic approach based on the continuous logarithm in the complex circle, designed to address the challenges posed by quantum computing. By leveraging its multi-valued and spectral properties, this framework enables the reintroduction of classical algorithms (DH, ECDSA, ElGamal, EC) and elliptic curve variants into the pos

  50. Jing Bi, Junjia Guo, Yunlong Tang, Lianggong Bruce Wen

    Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable progress in visual understanding. This impressive leap raises a compelling question: how can language models, initially trained solely on linguistic data, effectively interpret and process visual content? This paper aims to address this question with systematic inves

  51. Jiaxing Yu, Xinda Wu, Yunfei Xu, Tieyao Zhang

    Lyric-to-melody generation aims to automatically create melodies based on given lyrics, requiring the capture of complex and subtle correlations between them. However, previous works usually suffer from two main challenges: 1) lyric-melody alignment modeling, which is often simplified to one-syllable/word-to-one-note alignment, while others have the problem

  52. Mingcong Song, Xinru Tang, Fengfan Hou, Jing Li

    Meeting growing demands for low latency and cost efficiency in production-grade large language model (LLM) serving systems requires integrating advanced optimization techniques. However, dynamic and unpredictable input-output lengths of LLM, compounded by these optimizations, exacerbate the issues of workload variability, making it difficult to maintain high

  53. Lucas Fernando Alvarenga e Silva, Samuel Felipe dos Santos, Nicu Sebe, Jurandy Almeida

    Convolutional neural networks (CNNs) can learn directly from raw data, resulting in exceptional performance across various research areas. However, factors present in non-controllable environments such as unlabeled datasets with varying levels of domain and category shift can reduce model accuracy. The Open Set Domain Adaptation (OSDA) is a challenging probl

  54. Jason Langley, Kitzia Solis, Vala Masjedizadeh, Murphy Shao

    Elevated kurtosis values have been observed in subcortical grey matter structures of patients with neurodegenerative diseases. Here we tested whether these elevated values are related to iron content, which generate magnetic fields that add to the diffusion encoding gradients. Multi-shell diffusion and multi-echo gradient echo acquisitions were used to deriv

  55. Zhaomeng Deng, Ziqi Zhang, Ding Li, Yao Guo

    Real-time operating systems employ spatial and temporal isolation to guarantee predictability and schedulability of real-time systems on multi-core processors. Any unbounded and uncontrolled cross-core performance interference poses a significant threat to system time safety. However, the current Linux kernel has a number of interference issues and represent

  56. Yan Jiang, Xiaoyu Ji, Yancheng Jiang, Kai Wang

    Sensors are key components enabling various applications, e.g., home intrusion detection and environmental monitoring. While various software defenses and physical protections are used to prevent sensor manipulation, this paper introduces a new threat vector, PowerRadio, that bypasses existing protections and changes sensor readings from a distance. PowerRad

  57. Laiyuan Gao, Shicheng Zhang, Yuntao Zhang

    Mayer asks a question what closed, embedded and nonconvex initial curves guarantee that Gage's area-preserving flow (GAPF) exists globally. A folklore conjecture since 2012 says that GAPF evolves smooth, embedded and star-shaped initial curves globally. In this paper, we prove this conjecture by using Dittberner's singularity analysis theory. A star-shaped `

  58. Naota Sekiguchi, Yuta Kainuma, Motofumi Fushimi, Chikara Shinei

    We employ a dry-type phantom to evaluate the performance of a diamond quantum magnetometer with a high sensitivity of about $6~\mathrm{pT/\sqrt{Hz}}$ from the viewpoint of practical measurement in biomagnetic sensing. The dry phantom is supposed to represent an equivalent current dipole (ECD) generated by brain activity, emulating an encephalomagnetic field.

  59. Suyuan Wang, Xueqian Yin, Menghao Wang, Ruofeng Guo

    The rapid growth of scientific techniques and knowledge is reflected in the exponential increase in new patents filed annually. While these patents drive innovation, they also present significant burden for researchers and engineers, especially newcomers. To avoid the tedious work of navigating a vast and complex landscape to identify trends and breakthrough

  60. Yu-Ming Huang, Kuan-Yu Chen, Wen-Wei Lin, Da-Yi Chen

    Earthquake early warning systems play crucial roles in reducing the risk of seismic disasters. Previously, the dominant modeling system was the single-station models. Such models digest signal data received at a given station and predict earth-quake parameters, such as the p-phase arrival time, intensity, and magnitude at that location. Various methods have

  61. Tirawut Worrakitpoonpon

    We investigate the bar formation process using $N$-body simulations across the Toomre's parameter $Q_{min}$ and central mass concentration (CMC), focusing principally on the formation timescale. Of importance is that, as suggested by cosmological simulations, disk galaxies have limited time of $\sim 8$ Gyr in the Universe timeline to evolve secularly, starti

  62. Nan Yang, Chong Wang, Meihua Zhao, Zimeng Zhao

    Ocean forecasting is crucial for both scientific research and societal benefits. Currently, the most accurate forecasting systems are global ocean forecasting systems (GOFSs), which represent the ocean state variables (OSVs) as discrete grids and solve partial differential equations (PDEs) governing the transitions of oceanic state variables using numerical

  63. Yu He Ke, Liyuan Jin, Kabilan Elangovan, Bryan Wen Xi Ong

    Large Language Models (LLMs) are emerging as powerful tools in healthcare, particularly for complex, domain-specific tasks. This study describes the development and evaluation of the PErioperative AI CHatbot (PEACH), a secure LLM-based system integrated with local perioperative guidelines to support preoperative clinical decision-making. PEACH was embedded w

  64. Takaaki Nomura, Hiroshi Okada

    We study a Zee model in a non-holomorphic modular $A_4$ flavor symmetry in which we find good predictions in both the cases of normal and inverted hierarchy. Parameter reduction on neutrino sector occurs due to large mass hierarchies between charged-leptons mass eigenvalues and new singly-charged bosons in addition to this flavor symmetry. As a result, we ha

  65. Shyam Kumar M, Jiarong Hong

    Advanced three-dimensional (3D) tracking methods are essential for studying particle dynamics across a wide range of complex systems, including multiphase flows, environmental and atmospheric sciences, colloidal science, biological and medical research, and industrial manufacturing processes. This review provides a comprehensive summary of 3D particle tracki

  66. Rui Xiao, Jiong Wang, Lu Han, Na Zong

    Applying large language models (LLMs) as teaching assists has attracted much attention as an integral part of intelligent education, particularly in computing courses. To reduce the gap between the LLMs and the computer programming education expert, fine-tuning and retrieval augmented generation (RAG) are the two mainstream methods in existing researches. Ho

  67. Tuan-Nghia Bui, Huy-Son Nguyen, Cam-Van Nguyen Thi, Hoang-Quynh Le

    Bundle recommendation aims to suggest a set of interconnected items to users. However, diverse interaction types and sparse interaction matrices often pose challenges for previous approaches in accurately predicting user-bundle adoptions. Inspired by the distant supervision strategy and generative paradigm, we propose BRIDGE, a novel framework for bundle rec

  68. Lixian Jing, Jianpeng Qi, Junyu Dong, Yanwei Yu

    As deep neural networks (DNNs) are increasingly deployed on edge devices, optimizing models for constrained computational resources is critical. Existing auto-pruning methods face challenges due to the diversity of DNN models, various operators (e.g., filters), and the difficulty in balancing pruning granularity with model accuracy. To address these limitati

  69. Kanoko Goto, Takumi Karasawa, Takumi Hirose, Rei Kawakami

    Small object detection aims to localize and classify small objects within images. With recent advances in large-scale vision-language pretraining, finetuning pretrained object detection models has emerged as a promising approach. However, finetuning large models is computationally and memory expensive. To address this issue, this paper introduces multi-point

  70. Qijie Wei, Weihong Yu, Xirong Li

    Previous research on retinal vessel segmentation is targeted at a specific image domain, mostly color fundus photography (CFP). In this paper we make a brave attempt to attack a more challenging task of broad-domain retinal vessel segmentation (BD-RVS), which is to develop a unified model applicable to varied domains including CFP, SLO, UWF, OCTA and FFA. To

  71. Marius Tărnăuceanu

    Let $Ab_0$ be the class of finite abelian groups and consider the function $f:Ab_0\longrightarrow(0,\infty)$ given by $f(G)=\frac{|{\rm Aut}(G)|}{|G|}$\,, where ${\rm Aut}(G)$ is the automorphism group of a finite abelian group $G$. In this short note, we prove that the image of $f$ is a dense set in $[0,\infty)$.

  72. Dongbin Shin, Fabijan Pavošević, Nicolas Tancogne-Dejean, Michele Buzzi

    Recent studies of organic molecular solids are highlighted by their complex phase diagram and light-induced phenomena, such as Mott insulator, spin liquid phase, and superconductivity. However, a discrepancy between experimental observation and first-principle calculation on the $\kappa$-(BEDT-TTF)$_2$X family inhibits understanding their properties. Here, w

  73. Marius Tărnăuceanu

    T.C. Burness and S.D. Scott \cite{3} classified finite groups $G$ such that the number of prime order subgroups of $G$ is greater than $|G|/2-1$. In this note, we study finite groups $G$ whose subgroup graph contains a vertex of degree greater than $|G|/2-1$. The classification given for finite solvable groups extends the work of Burness and Scott.

  74. Aizierjiang Aiersilan

    Motion planning is a crucial component in autonomous driving. State-of-the-art motion planners are trained on meticulously curated datasets, which are not only expensive to annotate but also insufficient in capturing rarely seen critical scenarios. Failing to account for such scenarios poses a significant risk to motion planners and may lead to incidents dur

  75. Alexandr Garbali, Weiying Guo, Michael Wheeler

    Starting from the Izergin-Korepin 19-vertex model in the quadrant, we introduce two families of rational multivariate functions $F_S$ and $G_S$; these are in direct analogy with functions introduced by Borodin in the context of the higher-spin 6-vertex model in the quadrant. We prove that $F_S(x_1,\dots,x_N;z)$ and $G_S(y_1,\dots,y_M;z)$ are symmetric functi

  76. Mengke, Ma, Zilin Bian, Jingqin Gao

    Transportation equity research has traditionally emphasized service accessibility and destination reachability while often overlooking the critical aspects of service quality, such as infrequent schedules or overcrowded vehicles. This oversight can lead to a skewed understanding of equity, as high accessibility does not guarantee high-quality service. Addres

  77. Wentao Liu, Di Wu, Jieci Wang

    Among the three known types of static solutions proposed within the Hamiltonian constraint approach to effective quantum gravity (EQG), the first two have been extensively investigated, whereas the third type-which preserves general covariance, is free of Cauchy horizons, and was only recently obtained-remains relatively unexplored. This solution can describ

  78. Yuezihan Jiang, Gaode Chen, Wenhan Zhang, Jingchi Wang

    The item cold-start problem is crucial for online recommender systems, as the success of the cold-start phase determines whether items can transition into popular ones. Prompt learning, a powerful technique used in natural language processing (NLP) to address zero- or few-shot problems, has been adapted for recommender systems to tackle similar challenges. H

  79. Jae Ho Chang, Massimiliano Russo, Subhadeep Paul

    We study Heterogeneous Transfer Learning (HTL) for high-dimensional regression with differing feature sets. Such feature mismatch arises when some variables available in a data-rich source domain are unavailable in a data-poor target domain. Yet most homogeneous TL methods require the same feature space in both the source and target domains, limiting their p

  80. Victor Chernozhukov, Whitney K. Newey, Vasilis Syrgkanis

    There are many nonparametric objects of interest that are a function of a conditional distribution. One important example is an average treatment effect conditional on a subset of covariates. Many of these objects have a conditional influence function that generalizes the classical influence function of a functional of a (unconditional) distribution. Conditi

  81. Cristina Benea, Itamar Oliveira

    We prove the boundedness of a trilinear operator that is modulation invariant and which contains curvature information given by the presence of a complex exponential, adding to the small class of examples of such operators.

  82. Johannes Janning, Sophie F. Armanini, Urban Fasel

    The rapid development of advanced urban air mobility, particularly electric vertical take-off and landing (eVTOL) aircraft, requires interdisciplinary approaches involving the future urban air mobility ecosystem. Operational cost efficiency, regulatory aspects, sustainability, and environmental compatibility should be incorporated directly into the conceptua

  83. Chengwu Huang, U-Wai Lok, Jingke Zhang, Xiang Yang Zhu

    Ultrasound localization microscopy (ULM) enables microvascular imaging at spatial resolutions beyond the acoustic diffraction limit, offering significant clinical potentials. However, ULM performance relies heavily on microbubble (MB) signal sparsity, the number of detected MBs, and signal-to-noise ratio (SNR), all of which vary in clinical scenarios involvi

  84. Chang Liu, Xin Ma, Xiaochen Yang, Yuxiang Zhang

    Single-modal object detection tasks often experience performance degradation when encountering diverse scenarios. In contrast, multimodal object detection tasks can offer more comprehensive information about object features by integrating data from various modalities. Current multimodal object detection methods generally use various fusion techniques, includ

  85. Anthony Bosman, Christopher William Davis, Taylor Martin, Carolyn Otto

    The homotopy trivializing number, \(n_h(L)\), and the Delta homotopy trivializing number, \(n_\Delta(L)\), are invariants of the link homotopy class of \(L\) which count how many crossing changes or Delta moves are needed to reduce that link to a homotopy trivial link. In 2022, Davis, Orson, and Park proved that the homotopy trivializing number of \(L\) is b

  86. Fei Wu, Thomas Thiery, Stefanos Leonardos, Carmine Ventre

    Block production on the Ethereum blockchain has adopted an auction-based mechanism known as Proposer--Builder Separation (PBS), where validators outsource block creation to builders competing in MEV--Boost auctions for Maximal Extractable Value (MEV) rewards. We employ empirical game-theoretic analysis based on simulations to examine how advantages in latenc

  87. Yizhou Zhang, Yang Sui

    This paper explores the intricate behavior of deep neural networks (DNNs) through the lens of neuron activation dynamics. We propose a probabilistic framework that can analyze models' neuron activation patterns as a stochastic process, uncovering theoretical insights into neural scaling laws, such as over-parameterization and the power-law decay of loss with

  88. Wan-Cyuan Fan, Tanzila Rahman, Leonid Sigal

    With advances in foundational and vision-language models, and effective fine-tuning techniques, a large number of both general and special-purpose models have been developed for a variety of visual tasks. Despite the flexibility and accessibility of these models, no single model is able to handle all tasks and/or applications that may be envisioned by potent

  89. Christopher Kuo, Harold Williams

    We study the coamoebae of Lagrangian submanifolds of $(\mathbb{C}^\times)^n$, specifically how the combinatorics of their degenerations encodes the homological algebra of mirror coherent sheaves. Concretely, to a minimal free resolution $F^\bullet$ of a module $M$ over $\mathbb{C}[z_1^{\pm 1}, \dotsc, z_n^{\pm 1}]$ we associate a simplicial complex $T(F^\bul

  90. Ewan Davies, Olivia LeBlanc

    We study the maximum and minimum occupancy fraction of the antiferromagnetic Ising model in regular graphs. The minimizing problem is known to determine a computational threshold in the complexity of approximately sampling from the Ising model at a given magnetization, and our results determine this threshold for nearly the entire relevant parameter range in

  91. Mingda Chen, Yang Li, Karthik Padthe, Rulin Shao

    Large language models can generate factually inaccurate content, a problem known as hallucination. Recent works have built upon retrieved-augmented generation to improve factuality through iterative prompting but these methods are limited by the traditional RAG design. To address these challenges, we introduce EWE (Explicit Working Memory), a novel approach

  92. Joanna Bławat, Grzegorz Chajewski, Daniel Gnida, John Singleton

    CeRh2As2 is rare among superconductors, in that magnetic field tunes it between two distinct superconducting phases. Combined with a lack of local inversion symmetry and an upper critical field exceeding the Pauli paramagnetic limit, this excitingly suggests triplet multicomponent superconductivity. Preceding the superconducting onset, f-electron correlation

  93. Yu Liu, Aditya Raghavan, Utkarsh Pratiush, Maxim Ziatdinov

    Combinatorial materials libraries provide a powerful platform for mapping how physical properties evolve across binary and ternary cross-sections of multicomponent phase diagrams. While synthesis of such libraries has advanced since the 1960s and been accelerated by laboratory automation, their broader utility depends on rapid, quantitative measurements of c

  94. Marcel Valovy, Alena Buchalcevova

    This study aims to integrate blockchain technology into personality-based pair programming research to enhance its generalizability and adaptability by offering built-in continuous, reproducible, and transparent research. In the developing Role-Optimization Motivation Alignment (ROMA) framework, human/AI programming roles align with individual Big Five perso

  95. Yingjie Ma, Zitong Yu, Xun Lin, Weicheng Xie

    In the domain of facial recognition security, multimodal Face Anti-Spoofing (FAS) is essential for countering presentation attacks. However, existing technologies encounter challenges due to modality biases and imbalances, as well as domain shifts. Our research introduces a Mixture of Experts (MoE) model to address these issues effectively. We identified thr

  96. Junyan Zhao

    The moduli space of bundle stable pairs $\overline{M}_C(2,\Lambda)$ on a smooth projective curve $C$, introduced by Thaddeus, is a smooth Fano variety of Picard rank two. Focusing on the genus two case, we show that its K-moduli space is isomorphic to a GIT moduli of lines in quartic del Pezzo threefolds. Additionally, we construct a natural forgetful morphi

  97. Osama Hosam Abdellaif, Abdelrahman Nader, Ali Hamdi

    This paper introduces LMRPA, a novel Large Model-Driven Robotic Process Automation (RPA) model designed to greatly improve the efficiency and speed of Optical Character Recognition (OCR) tasks. Traditional RPA platforms often suffer from performance bottlenecks when handling high-volume repetitive processes like OCR, leading to a less efficient and more time

  98. Askold Khovanskii, Aaron Tronsgard

    We consider the problem of solvability of linear differential equations over a differential field~$K$. We introduce a class of special differential field extensions, which widely generalizes the classical class of extensions of differential fields by integrals and by exponentials of integrals and which has similar properties. We announce the following result

  99. Hyunbae Jeon, Frederic Guintu, Rayvant Sahni

    Turn-taking prediction is the task of anticipating when the speaker in a conversation will yield their turn to another speaker to begin speaking. This project expands on existing strategies for turn-taking prediction by employing a multi-modal ensemble approach that integrates large language models (LLMs) and voice activity projection (VAP) models. By combin

  100. Wen Wen, Yilin Wang, Neil Birkbeck, Balu Adsumilli

    The rise of short-form videos, characterized by diverse content, editing styles, and artifacts, poses substantial challenges for learning-based blind video quality assessment (BVQA) models. Multimodal large language models (MLLMs), renowned for their superior generalization capabilities, present a promising solution. This paper focuses on effectively leverag