Skip to content

March 2025 arXiv papers — page 53

Showing 5,2015,300 of 23,633 papers

  1. Mengming Li, Qijun Zhang, Yongqing Ren, Zhiyao Xie

    Hardware prefetching plays a critical role in hiding the off-chip DRAM latency. The complexity of applications results in a wide variety of memory access patterns, prompting the development of numerous cache-prefetching algorithms. Consequently, commercial processors often employ a hybrid of these algorithms to enhance the overall prefetching performance. No

  2. Zahra Hamed-Labbafian, Narjes Sabeghi, Mostafa Tavakoli, Sandi Klavžar

    If $G$ is a graph, then $X\subseteq V(G)$ is a general position set if for every two vertices $v,u\in X$ and every shortest $(u,v)$-path $P$, it holds that no inner vertex of $P$ lies in $X$. In this note we propose three algorithms to compute a largest general position set in $G$: an integer linear programming algorithm, a genetic algorithm, and a simulated

  3. Paul X. McCarthy, Xian Gong, Marieth Coetzer, Marian-Andrei Rizoiu

    This study explores the relationship between personality diversity and national economic performance, introducing the Global Personality Diversity Index ($Ψ$-GPDI) as a novel metric. Leveraging a dataset of 760,242 individuals across 135 countries, we quantify within-country diversity based on the Big Five personality traits. Our findings reveal that persona

  4. Yonatan Blumenthal, Uriya First

    Let $F$ be a field. We show that the largest irredundant generating sets for the algebra of $n\times n $ matrices over $F$ have $2n-1$ elements when $n>1$. (A result of Laffey states that the answer is $2n-2$ when $n>2$, but its proof contains an error.) We further give a classification of the largest irredundant generating sets when $n\in\{2,3\}$ and $F$ is

  5. Peishan Huang, Dong Li

    In recent years, the rapid development of machine learning has brought reforms and challenges to traditional communication systems. Semantic communication has appeared as an effective strategy to effectively extract relevant semantic signals semantic segmentation labels and image features for image transmission. However, the insufficient number of extracted

  6. Xiaohe Li, Haohua Wu, Jiahao Li, Zide Fan

    The rapid increase in remote sensing satellites has led to the emergence of distributed space-based observation systems. However, existing distributed remote sensing models often rely on centralized training, resulting in data leakage, communication overhead, and reduced accuracy due to data distribution discrepancies across platforms. To address these chall

  7. Jaihoon Kim, Taehoon Yoon, Jisung Hwang, Minhyuk Sung

    We propose an inference-time scaling approach for pretrained flow models. Recently, inference-time scaling has gained significant attention in LLMs and diffusion models, improving sample quality or better aligning outputs with user preferences by leveraging additional computation. For diffusion models, particle sampling has allowed more efficient scaling due

  8. Dimitri Loutchko, Yuki Sughiyama, Tetsuya J. Kobayashi

    Information geometry is based on classical Legendre duality but allows to incorporate additional structure such as algebraic constraints and Bregman divergence functions. It is naturally suited, and has been successfully used, to describe the thermodynamics of chemical reaction networks (CRNs) based on the Legendre duality between concentration and potential

  9. Yuan Wei, Xiaohan Shan, Jianmin Li

    Multi-agent reinforcement learning (MARL) faces two critical bottlenecks distinct from single-agent RL: credit assignment in cooperative tasks and partial observability of environmental states. We propose LERO, a framework integrating Large language models (LLMs) with evolutionary optimization to address these MARL-specific challenges. The solution centers o

  10. Yukang Lin, Hokit Fung, Jianjin Xu, Zeping Ren

    Recent portrait animation methods have made significant strides in generating realistic lip synchronization. However, they often lack explicit control over head movements and facial expressions, and cannot produce videos from multiple viewpoints, resulting in less controllable and expressive animations. Moreover, text-guided portrait animation remains undere

  11. Yuhan Wang, Silu He, Qinyao Luo, Hongyuan Yuan

    The existing methods learn geographic network representations through deep graph neural networks (GNNs) based on the i.i.d. assumption. However, the spatial heterogeneity and temporal dynamics of geographic data make the out-of-distribution (OOD) generalisation problem particularly salient. The latter are particularly sensitive to distribution shifts (featur

  12. Henri Aïdasso, Francis Bordeleau, Ali Tizghadam

    Despite the indisputable benefits of Continuous Integration (CI) pipelines (or builds), CI still presents significant challenges regarding long durations, failures, and flakiness. Prior studies addressed CI challenges in isolation, yet these issues are interrelated and require a holistic approach for effective optimization. To bridge this gap, this paper pro

  13. Yiwei Zhang

    This study proposes a risk pricing anomaly detection method for social network user portraits based on graph neural networks (GNNs), aiming to improve the ability to identify abnormal users in social network environments. In view of the limitations of traditional methods in social network data modeling, this paper combines graph autoencoders (GAEs) and graph

  14. Chenhao Jin, Yinhua Xia, Yan Xu

    We present a kernel compensation method for Maxwell eigenproblem for photonic crystals to avoid the infinite-dimensional kernels that cause many difficulties in the calculation of energy gaps. The quasi-periodic problem is first transformed into a periodic one on the cube by the Floquet-Bloch theory. Then the compensation operator is introduced in Maxwell's

  15. Sang-Hyun Chin, Daseul Lee, Donggyu Lee, Kwanghyun Chung

    Metal-organic chalcogenides (MOCs), robust crystalline assemblies composed of coinage metals, chalcogens and organic ligands, are typically synthesized via prolonged, high temperature tarnishing of vacuum-deposited metal films with organochalcogen precursors. The prolonged exposure to high temperatures and the necessity for direct vacuum deposition of silver

  16. Akshay Kulkarni, Ge Yan, Chung-En Sun, Tuomas Oikarinen

    Concept bottleneck models (CBM) aim to produce inherently interpretable models that rely on human-understandable concepts for their predictions. However, existing approaches to design interpretable generative models based on CBMs are not yet efficient and scalable, as they require expensive generative model training from scratch as well as real images with l

  17. Bijay Kumar Sahoo, Abhiram Soori

    We study a multi-terminal Josephson junction consisting of a central spin-orbit-coupled (SOC) region with an in-plane Zeeman field connected to four superconducting terminals. This setup allows for the simultaneous measurement of both longitudinal and transverse Josephson currents in response to a phase bias and provides a platform to probe the planar Hall e

  18. Dominic K Devlin, Austen RD Ganley, Nobuto Takeuchi

    Morphogenesis of complex body shapes is reproducible despite the noise inherent in the underlying morphogenetic processes. However, how these morphogenetic processes work together to achieve this reproducibility remains unclear. Here, we ask how morphogenetic reproducibility is realised by developing a computational model that evolves complex morphologies. W

  19. Ying Zhang, Atsushi Hosaka, Qian Wang, Shigehiro Yasui

    Studying exotic hadrons is a challenge against the conventional quark model, providing us with a good platform to deepen our understanding of the strong interaction. An inclusive study of the exotic hadrons in vacuum and at finite temperature is an intriguing approach to shed light on their nature. As a first step, we study the $Z_c(3900)$ in both the $D\bar

  20. Hyeongjin Nam, Donghwan Kim, Jeongtaek Oh, Kyoung Mu Lee

    Most existing methods of 3D clothed human reconstruction from a single image treat the clothed human as a single object without distinguishing between cloth and human body. In this regard, we present DeClotH, which separately reconstructs 3D cloth and human body from a single image. This task remains largely unexplored due to the extreme occlusion between cl

  21. Vladimir P. Savin, Yury A. Koksharov

    The possible manifestations of magnetic screening effects were theoretically investigated for a particle with a ferromagnetic single-domain core and a magnetically soft shell. The exact solution of the Laplace equation gave analytical formulas for the magnetic field created by the particle placed in a uniform external magnetic field. The system of non-intera

  22. Heqing Zou, Fengmao Lv, Desheng Zheng, Eng Siong Chng

    Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses significant challenges for zero-shot speech emotion recognition, especially with multilingual datasets. In this paper, we propose leveraging co

  23. Daniel Saragih, Deyu Cao, Tejas Balaji, Ashwin Santhosh

    Foundational language models show a remarkable ability to learn new concepts during inference via context data. However, similar work for images lag behind. To address this challenge, we introduce FLoWN, a flow matching model that learns to generate neural network parameters for different tasks. Our approach models the flow on latent space, while conditionin

  24. Taishin Saito

    In order to understand the overall picture of cyber attacks and to identify the source of cyber attacks, a method to identify malicious activities by automatically creating a graph that ties together the dependencies of a series of related events by tracking Data Provenance has been developed. However, the problem of dependency explosion, in which a large nu

  25. Yufei Cai, Hu Han, Yuxiang Wei, Shiguang Shan

    The progress on generative models has led to significant advances on text-to-video (T2V) generation, yet the motion controllability of generated videos remains limited. Existing motion transfer methods explored the motion representations of reference videos to guide generation. Nevertheless, these methods typically rely on sample-specific optimization strate

  26. Jiawei Yao, Yijie Mao, Mingzhe Chen, Ye Hu

    Reconfigurable Intelligent Surface (RIS) has been recognized as a promising solution for enhancing localization accuracy. Traditional RIS-based localization methods typically rely on prior channel knowledge, beam scanning, and pilot-based assistance. These approaches often result in substantial energy and computational overhead, and require real-time coordin

  27. Zizhi Chen, Minghao Han, Xukun Zhang, Shuwei Ma

    Multimodal learning combining pathology images and genomic sequences enhances cancer survival analysis but faces clinical implementation barriers due to limited access to genomic sequencing in under-resourced regions. To enable survival prediction using only whole-slide images (WSI), we propose the Visual-Genomic Answering-Guided Transformer (VGAT), a framew

  28. Jiaxuan Wu, Wanli Peng, Hang Fu, Yiming Xue

    Training large language models (LLMs) is resource-intensive and expensive, making protecting intellectual property (IP) for LLMs crucial. Recently, embedding fingerprints into LLMs has emerged as a prevalent method for establishing model ownership. However, existing fingerprinting techniques typically embed identifiable patterns with weak semantic coherence,

  29. Ningning Tao, Xiaosong Chen, Fei Xie, Yongwen Zhang

    Variations in stratospheric atmospheric circulation significantly influence tropospheric weather and climate, and understanding these variations can guide stratospheric aircraft development and operations. Despite a century of progress, large-scale patterns in stratospheric circulation remain poorly understood due to the stratosphere's complex nature. To add

  30. Klaus Jansen, Debajyoti Kar, Arindam Khan, K. V. N. Sreenivas

    We study the three-dimensional Knapsack (3DK) problem, in which we are given a set of axis-aligned cuboids with associated profits and an axis-aligned cube knapsack. The objective is to find a non-overlapping axis-aligned packing (by translation) of the maximum profit subset of cuboids into the cube. The previous best approximation algorithm is due to Diedri

  31. Zhang Di, Wang Zeyin, Tang Yanqun, Wu Dongdong

    The secure affine frequency division multiplexing (AFDM) waveform design is a main concern in high-mobility networks. In this article, we employ the four key parameters in AFDM to design secure waveforms, and afterward we analyze the role of the four parameters to reveal the design guideline. We find that c1 is bounded by the Doppler shifts and preset guard.

  32. Nipen Saikia, Adam Paksok

    Alanzi et al. (2022) investigated overpartition of a positive integer $n$ with $\ell$-regular non-overlined parts denoted by $\overline R_\ell^\ast (n)$, and proved some results for the case $\ell=3$. As extension to the results of Alanzi et al., Sellers (2024) proved some new congruences for $\overline R_3^\ast (n)$. In this paper, we prove some new infinit

  33. Huaiqian Li, Bingyao Wu

    We investigate the limiting behavior of Besov seminorms and nonlocal perimeters in Dunkl theory. The present work generalizes two fundamental results: the Maz'ya--Shaposhnikova formula for Gagliardo seminorms and the asymptotics of (relative) fractional $s$-perimeters. Our main contributions are twofold. First, we establish a dimension-free Maz'ya--Shaposhni

  34. Ying-Jung Chen, Ahmad Albarqawi, Chi-Sheng Chen

    Recent advances in the data-driven medicine approach, which integrates ethically managed and explainable artificial intelligence into clinical decision support systems (CDSS), are critical to ensure reliable and effective patient care. This paper focuses on comparing novel agent system designs that use modular agents to analyze laboratory results, vital sign

  35. Sujan Kumar Roy, Gargi Chaudhuri

    A number of hadronic equations of state for neutron stars have been investigated for the purpose of the present paper, considering the fact that at sufficiently high density, heavy baryons and quark phases may appear. The observational limits from NICER, GW170817, etc., are obeyed by our choice of equations of state. The universal relations are investigated

  36. Piera Riccio, Francesco Galati, Kajetan Schweighofer, Noa Garcia

    In the era of large-scale visual data, understanding collections of images is a challenging yet important task. To this end, we introduce ImageSet2Text, a novel method to automatically generate natural language descriptions of image sets. Based on large language models, visual-question answering chains, an external lexical graph, and CLIP-based verification,

  37. Jiawen Zhang, Zhenqi Hua, Chengwei Wang, Michael Smidman

    Introducing the concept of topology into material science has sparked a revolution from classic electronic and optoelectronic devices to topological quantum devices. The latter has potential for transferring energy and information with unprecedented efficiency. Here, we demonstrate a topological diode effect on the surface of a three-dimensional material, Sm

  38. Yunhe Gao, Di Liu, Zhuowei Li, Yunsheng Li

    Medical image segmentation remains challenging due to the vast diversity of anatomical structures, imaging modalities, and segmentation tasks. While deep learning has made significant advances, current approaches struggle to generalize as they require task-specific training or fine-tuning on unseen classes. We present Iris, a novel In-context Reference Image

  39. Zhiwei Huang, Hailin Yu, Yichun Shentu, Jin Yuan

    This paper presents a novel camera relocalization method, STDLoc, which leverages Feature Gaussian as scene representation. STDLoc is a full relocalization pipeline that can achieve accurate relocalization without relying on any pose prior. Unlike previous coarse-to-fine localization methods that require image retrieval first and then feature matching, we pr

  40. Farzad Beizaee, Gregory A. Lodygensky, Christian Desrosiers, Jose Dolz

    Recent advances in diffusion models have spurred research into their application for Reconstruction-based unsupervised anomaly detection. However, these methods may struggle with maintaining structural integrity and recovering the anomaly-free content of abnormal regions, especially in multi-class scenarios. Furthermore, diffusion models are inherently desig

  41. Reza Pourreza, Rishit Dagli, Apratim Bhattacharyya, Sunny Panchal

    AI models have made significant strides in recent years in their ability to describe and answer questions about real-world images. They have also made progress in the ability to converse with users in real-time using audio input. This raises the question: have we reached the point where AI models, connected to a camera and microphone, can converse with users

  42. Dohwan Ko, Sihyeon Kim, Yumin Suh, Vijay Kumar B. G

    Spatio-temporal reasoning is essential in understanding real-world environments in various fields, eg, autonomous driving and sports analytics. Recent advances have improved the spatial reasoning ability of Vision-Language Models (VLMs) by introducing large-scale data, but these models still struggle to analyze kinematic elements like traveled distance and s

  43. Yuta Hirabayashi, Daisuke Matsuoka

    Data-driven weather prediction models exhibit promising performance and advance continuously. In particular, diffusion models represent fine-scale details without spatial smoothing, which is crucial for mesoscale predictions, such as heavy rainfall forecasting. However, the applications of diffusion models to mesoscale prediction remain limited. To address t

  44. Yuxuan Hu, Xiaodong Chen, Cuiping Li, Hong Chen

    Large Language Models (LLMs) excel in diverse applications but suffer inefficiency due to massive scale. While quantization reduces computational costs, existing methods degrade accuracy in medium-sized LLMs (e.g., Llama-3-8B) due to activation outliers. To address this, we propose QUAD (Quantization with Activation Decomposition), a framework leveraging Sin

  45. Tomoaki Ishiyama, Francisco Prada, Anatoly A. Klypin

    Observations favor cosmological models with a time-varying dark energy component. But how does dynamical dark energy (DDE) influence the growth of structure in an expanding Universe? We investigate this question using high-resolution $N$-body simulations based on a DDE cosmology constrained by first-year DESI data (DESIY1$+$DDE), characterized by a 4% lower

  46. Jingyu Liu, Zijie Xin, Yuhan Fu, Ruixiang Zhao

    Sketch animation, which brings static sketches to life by generating dynamic video sequences, has found widespread applications in GIF design, cartoon production, and daily entertainment. While current methods for sketch animation perform well in single-object sketch animation, they struggle in multi-object scenarios. By analyzing their failures, we identify

  47. Jie Gu, Yunfeng Jiang, Huajia Wang

    We elaborate on the resurgence analysis on the $T\overline{T}$-deformed 2d conformal field theory (CFT). Writing the deformed partition function as an infinite series in the deformation parameter $\lambda$, we develop efficient analytical methods to compute high-order terms of the $\lambda$-series. Based on the asymptotic behavior of the large-order perturba

  48. Shengbo Wang, Ke Li, Zheng Yan, Zhenyuan Guo

    Safety is of paramount importance in control systems to avoid costly risks and catastrophic damages. The control barrier function (CBF) method, a promising solution for safety-critical control, poses a new challenge of enhancing control performance due to its direct modification of original control design and the introduction of uncalibrated parameters. In t

  49. Amir Jalili, Feng Pan, Ai Xi Chen, Jerry P. Draayer

    This study investigates the application of deep learning models-recurrent neural networks, gated recurrent units, and long short-term memory networks-for predicting nuclear binding energies. Utilizing data from the Atomic Mass Evaluation (AME2020), we incorporate key nuclear structure features, including proton and neutron numbers, as well as additional term

  50. Philip Doldo, Derek Everett, Amol Khanna, Andre T Nguyen

    Projected Gradient Descent (PGD) under the $L_\infty$ ball has become one of the defacto methods used in adversarial robustness evaluation for computer vision (CV) due to its reliability and efficacy, making a strong and easy-to-implement iterative baseline. However, PGD is computationally demanding to apply, especially when using thousands of iterations is

  51. Jianbo Cui, Georg Maierhofer

    We introduce a novel approach to numerical approximation of nonlinear Schr\"odinger equation with white noise dispersion in the regime of low-regularity solutions. Approximating such solutions in the stochastic setting is particularly challenging due to randomized frequency interactions and presents a compelling challenge for the construction of tailored sch

  52. Dilip Kumar, Bijita Bose, Soma Sanyal

    Magnetic reconnection in magnetized wakes of cosmic strings results in the release of a large amount of energy. This energy is released in a short period of time. In this work, we show that this sudden release of energy can result in a Gamma Ray Burst (GRB) of short duration. The magnetic reconnection occurs at several points of the cosmic string wake. The e

  53. Shusaku Egami, Kyoumoto Matsushita, Takanori Ugai, Ken Fukuda

    Hyper-relational Knowledge Graphs (HRKGs) extend traditional KGs beyond binary relations, enabling the representation of contextual, provenance, and temporal information in domains, such as historical events, sensor data, video content, and narratives. HRKGs can be structured using several Metadata Representation Models (MRMs), including Reification (REF), S

  54. Foster Tom, Aarush Vailaya

    We describe how the chromatic symmetric function of two graphs glued at a single vertex can be expressed as a matrix multiplication using certain information of the two individual graphs. We then prove new $e$-positivity results by using a connection between forest triples, defined by the first author, and Hikita's probabilities associated to standard Young

  55. Sohan Kumar Jha

    In this article, we obtain a novel black hole (BH) solution of a Schwarzschild BH immersed in a Hernquist dark matter (SBHD) halo. The thermodynamic properties of the resultant spacetime are then studied to gauge the impact of dark matter (DM) on the local and global stability of the composite system of the BH-DM halo. With the intention of finding imprints

  56. Keerthy Menon, Thomas Busch, Thomás Fogarty

    A key focus of designing quantum thermal devices is the potential advantage that can be gleaned from genuine quantum effects when compared to classical devices. The recent experimental realization of the Pauli engine, where energy is extracted via changes in particle statistics as an alternative to conventional heat sources has opened new avenues of research

  57. Yinchuan Li, Guangchen Lan, Xiaodong Wang

    We propose a tensor generalized approximate message passing (TeG-AMP) algorithm for low-rank tensor inference, which can be used to solve tensor completion and decomposition problems. We derive TeG-AMP algorithm as an approximation of the sum-product belief propagation algorithm in high dimensions where the central limit theorem and Taylor series approximati

  58. Snehamoy Chatterjee, Greg Waite, Sidike Paheding, Luke Bowman

    Forecasting volcanic activity is critical for hazard assessment and risk mitigation. Volcanic Radiative Power (VPR), derived from thermal remote sensing data, serves as an essential indicator of volcanic activity. In this study, we employ Bayesian Regularized Neural Networks (BRNN) to predict future VPR values based on historical data from Fuego Volcano, com

  59. Yuguang Li, Ivaylo Boyadzhiev, Zixuan Liu, Linda Shapiro

    Reconstructing precise camera poses and floor plan layouts from wide-baseline RGB panoramas is a difficult and unsolved problem. We introduce BADGR, a novel diffusion model that jointly performs reconstruction and bundle adjustment (BA) to refine poses and layouts from a coarse state, using 1D floor boundary predictions from dozens of images of varying input

  60. Amna Naeem, Jawad Ahmad, Muazzam A. Khan, Aizaz Ahmad Khattak

    The ever-increasing security vulnerabilities in the Internet-of-Things (IoT) systems require improved threat detection approaches. This paper presents a compact and efficient approach to detect botnet attacks by employing an integrated approach that consists of traffic pattern analysis, temporal support learning, and focused feature extraction. The proposed

  61. Hengyu Wu, Yang Cao

    As large-scale models such as Large Language Models (LLMs) and Large Multimodal Models (LMMs) see increasing deployment, their privacy risks remain underexplored. Membership Inference Attacks (MIAs), which reveal whether a data point was used in training the target model, are an important technique for exposing or assessing privacy risks and have been shown

  62. Xiangji Cai, Yanyan Feng, Jing Ren, Kang Lan

    We theoretically study the quantum speed limits (QSLs) of a qubit system coupled to a thermal dephasing environment with an Ohmic-like spectral density. Based on the geometric QSLs time bound, which is derived by employing the trace distance to quantify the geodesic between two distinguishable states in dynamical evolution, we study the influences of the tem

  63. Guanxiong Qu

    Quantum geometry is a well-established framework for understanding transport and optical responses in quantum materials. In this work, I study the photon drag effect in Dirac electrons using the quantum geometric interpretation of non-vertical optical transitions. Due to the particle-hole symmetry inherent in Dirac electrons, the shift photon-drag photocurre

  64. Hoang Thieu Anh, Le Mau Hai, Nguyen Quang Dieu, Nguyen Van Phu

    In this paper, we study Hessian type equations for $(\o,m)-\beta$-subharmonic functions on a ball in $\mathbb{C}^n$, where $\beta=dd^c\|z\|^2=\frac{i}{2}\sum\limits_{j=1}^n dz_j\w d\bar{z}_j$ is the flat metric on $\cn$. Using the recent results in \cite{KN23b}, we are able to show the existence of bounded solutions for Hessian type equations.

  65. Ghazanfar Ali, Hong-Quan Le, Junho Kim, Seoung-won Hwang

    In this paper, we present the design of a multimodal interaction framework for intelligent virtual agents in wearable mixed reality environments, especially for interactive applications at museums, botanical gardens, and similar places. These places need engaging and no-repetitive digital content delivery to maximize user involvement. An intelligent virtual

  66. Bruno Jacob, Ashish S. Nair, Amanda A. Howard, Jan Drgona

    Physics-informed neural networks (PINNs) have demonstrated promise as a framework for solving forward and inverse problems involving partial differential equations. Despite recent progress in the field, it remains challenging to quantify uncertainty in these networks. While techniques such as Bayesian PINNs (B-PINNs) provide a principled approach to capturin

  67. Zhiying Yan, Yiyuan Liang, Shilv Cai, Tao Zhang

    Semantic 4D Gaussians can be used for reconstructing and understanding dynamic scenes, with temporal variations than static scenes. Directly applying static methods to understand dynamic scenes will fail to capture the temporal features. Few works focus on dynamic scene understanding based on Gaussian Splatting, since once the same update strategy is employe

  68. Chau Pham, Juan C. Caicedo, Bryan A. Plummer

    Prior work using Masked Autoencoders (MAEs) typically relies on random patch masking based on the assumption that images have significant redundancies across different channels, allowing for the reconstruction of masked content using cross-channel correlations. However, this assumption does not hold in Multi-Channel Imaging (MCI), where channels may provide

  69. Jee Won Lee, Hansol Lim, SooYeun Yang, Jongseong Brad Choi

    This paper presents a novel masked attention-based 3D Gaussian Splatting (3DGS) approach to enhance robotic perception and object detection in industrial and smart factory environments. U2-Net is employed for background removal to isolate target objects from raw images, thereby minimizing clutter and ensuring that the model processes only relevant data. Addi

  70. Yongting Hu, Yuxin Lin, Chengliang Liu, Xiaoling Luo

    Multi-view diabetic retinopathy (DR) detection has recently emerged as a promising method to address the issue of incomplete lesions faced by single-view DR. However, it is still challenging due to the variable sizes and scattered locations of lesions. Furthermore, existing multi-view DR methods typically merge multiple views without considering the correlat

  71. Vidya Srinivas, Xuhai Xu, Xin Liu, Kumar Ayush

    While NLP research has made strides in conversational tasks, many approaches focus on single-turn responses with well-defined objectives or evaluation criteria. In contrast, coaching presents unique challenges with initially undefined goals that evolve through multi-turn interactions, subjective evaluation criteria, mixed-initiative dialogue. In this work, w

  72. Sergio Torres Aguilar

    This paper introduces TRIDIS (Tria Digita Scribunt), an open-source corpus of medieval and early modern manuscripts. TRIDIS aggregates multiple legacy collections (all published under open licenses) and incorporates large metadata descriptions. While prior publications referenced some portions of this corpus, here we provide a unified overview with a stronge

  73. Zhen-Tai Zhang, Wei Zhong, Xianyu Tan, Bo Ma

    The scattering is crucial for the atmospheric thermal profiles. The energy transport by the vertical mixing plays an essential role for the greenhouse or anti-greenhouse effect. This work explores the interaction between scattering and vertical mixing, specifically whether these processes enhance or mitigate each other's effects on atmospheric temperature. T

  74. Yu Cui, Bryan Hooi, Yujun Cai, Yiwei Wang

    Recent reasoning large language models (LLMs) have demonstrated remarkable improvements in mathematical reasoning capabilities through long Chain-of-Thought. The reasoning tokens of these models enable self-correction within reasoning chains, enhancing robustness. This motivates our exploration: how vulnerable are reasoning LLMs to subtle errors in their inp

  75. Yuchao Gu, Weijia Mao, Mike Zheng Shou

    Long-context video modeling is essential for enabling generative models to function as world simulators, as they must maintain temporal coherence over extended time spans. However, most existing models are trained on short clips, limiting their ability to capture long-range dependencies, even with test-time extrapolation. While training directly on long vide

  76. Qi Li

    Center-based clustering algorithms (e.g., K-means) are popular for clustering tasks, but they usually struggle to achieve high accuracy on complex datasets. We believe the main reason is that traditional center-based clustering algorithms identify only one clustering center in each cluster. Once the distribution of the dataset is complex, a single clustering

  77. Trevor Karn, Victor Reiner

    This paper considers a finite group $G$ acting linearly on the variables $V$ of a polynomial algebra, or an exterior algebra, or superpolynomial algebra with both commuting and anticommuting variables. In this setting, the Hilbert series for the $G$-invariant subalgebra turns out to determine the analogous Hilbert series for the wreath product $P[G]$ acting

  78. Hiroyuki Tako Ishikawa, Stanimir Metchev, Megan E. Tannock, Gregory N. Mace

    We present a high signal-to-noise (SNR $\sim$ 450), high-dispersion ($R \equiv \lambda / \Delta \lambda \sim 28\,000$) H- and K-band spectroscopic atlas of the L7.5 and T0.5 components of the Luhman 16AB binary (WISE J104915.57$-$531906.1AB): the closest pair of brown dwarfs, and one of the best substellar benchmarks. The spectra were combined from a 70-day

  79. Ali Mohaghegh, Cheng Huang

    Though high-performance computing enables high-fidelity simulations of complex engineering systems, accurately resolving multi-scale physics for real-world problems remains computationally prohibitive, particularly in many-query applications such as optimization and uncertainty quantification. Projection-based model order reduction (MOR) has demonstrated sig

  80. Mahsa Paknejad, Parisa Fard Moshiri, Murat Simsek, Burak Kantarci

    This paper explores the advancement of Vehicular Edge Computing (VEC) as a tailored application of Mobile Edge Computing (MEC) for the automotive industry, addressing the rising demand for real-time processing in connected and autonomous vehicles. VEC brings computational resources closer to vehicles, reducing data processing delays crucial for safety-critic

  81. Parisa Fard Moshiri, Murat Simsek, Burak Kantarci

    The demand for MEC has increased with the rise of data-intensive applications and 5G networks, while conventional cloud models struggle to satisfy low-latency requirements. While task offloading is crucial for minimizing latency on resource-constrained User Equipment (UE), fully offloading of all tasks to MEC servers may result in overload and possible task

  82. Ahmed Omara, Burak Kantarci

    As Artificial Intelligence (AI) becomes increasingly integrated into microgrid control systems, the risk of malicious actors exploiting vulnerabilities in Machine Learning (ML) algorithms to disrupt power generation and distribution grows. Detection models to identify adversarial attacks need to meet the constraints of edge environments, where computational

  83. Shaoting Peng, Haonan Chen, Katherine Driggs-Campbell

    Learning human preferences is essential for human-robot interaction, as it enables robots to adapt their behaviors to align with human expectations and goals. However, the inherent uncertainties in both human behavior and robotic systems make preference learning a challenging task. While probabilistic robotics algorithms offer uncertainty quantification, the

  84. Zhiping Xiao, Xinyu Wang, Yifang Qin, Zijie Huang

    Understanding the evolution of public opinion is crucial for informed decision-making in various domains, particularly public affairs. The rapid growth of social networks, such as Twitter (now rebranded as X), provides an unprecedented opportunity to analyze public opinion at scale without relying on traditional surveys. With the rise of deep learning, Graph

  85. Xianshu Ju, Xiangkai Ke, Changhua Wei

    This paper is concerned with the global existence and blowup of the classical solution to the Cauchy problem of the relativistic Euler equation with $ p=0 $ in a fixed Friedmann-Lema\^{\i}tre-Robertson-Walker (FLRW) spacetime. The aim of this work is to study clearly the effect of the expansion rate of the spacetime on the life span of the classical solution

  86. Yuan Li, Jun Hu, Jiaxin Jiang, Zemin Liu

    Recent advances in graph learning have paved the way for innovative retrieval-augmented generation (RAG) systems that leverage the inherent relational structures in graph data. However, many existing approaches suffer from rigid, fixed settings and significant engineering overhead, limiting their adaptability and scalability. Additionally, the RAG community

  87. Chao Li, Boyu Zhang

    We construct stable minimal hypersurfaces with simple topology in certain compact $4$-manifolds $X$ with boundary, where $X$ embeds into a smooth manifold homeomorphic to $S^4$. For example, if $X$ is equipped with a Riemannian metric $g$ with positive scalar curvature, we prove the existence of a stable minimal hypersurface $M$ that is diffeomorphic to eith

  88. Jiaqi Liao, Zhengyuan Yang, Linjie Li, Dianqi Li

    In this work, we study the problem of Text-to-Image In-Context Learning (T2I-ICL). While Unified Multimodal LLMs (MLLMs) have advanced rapidly in recent years, they struggle with contextual reasoning in T2I-ICL scenarios. To address this limitation, we propose a novel framework that incorporates a thought process called ImageGen-CoT prior to image generation

  89. Weizhi Chen, Yupeng Deng, Jin Wei, Jingbo Chen

    Vision Language Foundation Models based on CLIP architecture for remote sensing primarily rely on short text captions, which often result in incomplete semantic representations. Although longer captions convey richer information, existing models struggle to process them effectively because of limited text-encoding capacity, and there remains a shortage of re

  90. Leonora Kaldaras, Carl Wieman

    Blended math science sensemaking (MSS) is reflected in a student ability to integrate math and science knowledge to develop mathematical descriptions of observations. While an important component of scientific thinking, there is little research on teaching MSS. Students from backgrounds historically marginalized in STEM often lack prior learning opportunitie

  91. Gollam Rabby, Diyana Muhammed, Prasenjit Mitra, Sören Auer

    Scientific hypothesis generation is a fundamentally challenging task in research, requiring the synthesis of novel and empirically grounded insights. Traditional approaches rely on human intuition and domain expertise, while purely large language model (LLM) based methods often struggle to produce hypotheses that are both innovative and reliable. To address

  92. Chaohan Wang, Yutong Xie, Qi Chen, Yuyin Zhou

    Mamba, with its selective State Space Models (SSMs), offers a more computationally efficient solution than Transformers for long-range dependency modeling. However, there is still a debate about its effectiveness in high-resolution 3D medical image segmentation. In this study, we present a comprehensive investigation into Mamba's capabilities in 3D medical i

  93. Zhuoran Zhao, Linlin Yang, Pengzhan Sun, Pan Hui

    Recent synthetic 3D human datasets for the face, body, and hands have pushed the limits on photorealism. Face recognition and body pose estimation have achieved state-of-the-art performance using synthetic training data alone, but for the hand, there is still a large synthetic-to-real gap. This paper presents the first systematic study of the synthetic-to-re

  94. Amjad Ali, Saeed Aldahmani, Hailiang Du, Zardad Khan

    This paper introduces the centroid decision forest (CDF), a novel ensemble learning framework that redefines the splitting strategy and tree building in the ordinary decision trees for high-dimensional classification. The splitting approach in CDF differs from the traditional decision trees in theat the class separability score (CSS) determines the selection

  95. Robin Cockett, Melika Norouzbeygi

    In this paper we prove that giving a right actegory with hom-objects is equivalent to giving a right-enriched category with copowers. While this result is known in the closed symmetric setting, our contribution extends the equivalence to non-closed and non-symmetric monoidal bases. This generalization is motivated by the semantics of higher-order message pas

  96. Yongxia Zhang, Jinwen Liang, Liwen Xu, Keming Yu

    This paper develops an inferential theory for high-dimensional matrix-variate factor models with missing observations. We propose an easy-to-use all-purpose method that involves two straightforward steps. First, we perform principal component analysis on two re-weighted covariance matrices to obtain the row and column loadings. Second, we utilize these loadi

  97. Hanshuo Qiu, Jie Jiang, Ruoli Yang, Lixin Zhan

    RGB-T road scene semantic segmentation enhances visual scene understanding in complex environments characterized by inadequate illumination or occlusion by fusing information from RGB and thermal images. Nevertheless, existing RGB-T semantic segmentation models typically depend on simple addition or concatenation strategies or ignore the differences between

  98. Yunuo Zhang, Baiting Luo, Ayan Mukhopadhyay, Abhishek Dubey

    Partially observable Markov decision processes (POMDPs) are a general mathematical model for sequential decision-making in stochastic environments under state uncertainty. POMDPs are often solved \textit{online}, which enables the algorithm to adapt to new information in real time. Online solvers typically use bootstrap particle filters based on importance r

  99. Naoki Yonezawa

    Blockchain consensus mechanisms must balance security, decentralization, and efficiency while ensuring fair participation. Proof of Team Sprint (PoTS) is a cooperative consensus mechanism designed to address the energy inefficiencies and centralization tendencies of traditional Proof of Work (PoW). Unlike PoW, where rewards disproportionately favor high-perf

  100. Narges Mehran, Nikolay Nikolov, Radu Prodan, Dumitru Roman

    The increased usage of Internet of Things devices at the network edge and the proliferation of microservice-based applications create new orchestration challenges in Edge computing. These include detecting overutilized resources and scaling out overloaded microservices in response to surging requests. This work presents ADApt, an extension of the ADA-PIPE to