Skip to content

October 2025 arXiv papers — page 183

Showing 18,20118,300 of 25,213 papers

  1. Junyu Shi, Minghui Li, Junguo Zuo, Zhifei Yu

    Deepfakes, leveraging advanced AIGC (Artificial Intelligence-Generated Content) techniques, create hyper-realistic synthetic images and videos of human faces, posing a significant threat to the authenticity of social media. While this real-world threat is increasingly prevalent, existing academic evaluations and benchmarks for detecting deepfake forgery ofte

  2. Deepika Gill, Sangeeta Sharma, Peter Elliott, Kay Dewhurst

    Graphene, and other members of the monolayer Xene family, represent an ideal materials platform for "valleytronics", the control of valley localized charge excitations. The absence of a gap in these semi-metals, however, precludes valley excitation by circularly polarized light pulses, sharply circumscribing the possibility of a lightwave valleytronics in th

  3. Qiuqi Li, Chang Liu, Yifei Yang

    Dynamic mode decomposition (DMD) is a widely used data-driven algorithm for predicting the future states of dynamical systems. However, its standard formulation often struggles with poor long-term predictive accuracy. To address this limitation, we propose a localized DMD (LDMD) framework that improves prediction performance by integrating DMD's strong linea

  4. Akashdeep Akashdeep, Sachin Krishnia, Jae-Hyun Ha, Siyeon An

    Ruthenium dioxide (RuO2) has recently emerged as a candidate altermagnet, yet its intrinsic magnetic ground state, particularly in thin films, remains debated. This study aims to clarify the nature and spatial extent of the magnetic order in RuO2 thin films grown under different conditions. Thin films of RuO2 with thicknesses of 30 nm and 33 nm are fabricate

  5. Fabio Morreale, Wiebke Hutiri, Joan Serrà, Alice Xiang

    The rise of AI-generated music is diluting royalty pools and revealing structural flaws in existing remuneration frameworks, challenging the well-established artist compensation systems in the music industry. Existing compensation solutions, such as piecemeal licensing agreements, lack scalability and technical rigour, while current data attribution mechanis

  6. Daniel Cohen Hillel

    A recent paper by Jordan et al. introduced Decoded Quantum Interferometry (DQI), a novel quantum algorithm that uses the quantum Fourier transform to reduce linear optimization problems -- max-XORSAT and max-LINSAT -- to decoding problems. In this paper, we extend DQI to optimization problems involving quadratic constraints, which we call max-QUADSAT. Levera

  7. Giulio Weikmann, Gianmarco Perantoni, Lorenzo Bruzzone

    This work presents a multitemporal class-driven hierarchical Residual Neural Network (ResNet) designed for modelling the classification of Time Series (TS) of multispectral images at different semantical class levels. The architecture consists of a modification of the ResNet where we introduce additional branches to perform the classification at the differen

  8. Timon Klein, Piotr Minakowski, Sebastian Sager, Steffen Schotthöfer

    Subject-specific distribution shifts represent a fundamental obstacle to developing foundation models for brain decoding. We propose the Subject-Specific Low-Rank Adapter (SuLoRA), a drop-in replacement for standard linear or convolutional layers that captures inter-subject variability by decomposing weights into a shared, subject-invariant component and a l

  9. Shule Lu, Lingxiang Wang, Sijia Wen, Ziwei Wang

    With the rapid development of artificial intelligence, dialogue systems have become a prominent form of human-computer interaction. However, traditional centralized or fully local training approaches face challenges in balancing privacy preservation and personalization due to data privacy concerns and heterogeneous device capabilities. Federated learning, as

  10. Hikmet Çakmak

    On 29 March 2006, a total solar eclipse was observed in the Manavgat district of Antalya, Turkey. During the event, the solar corona was observed using an 8-inch mirrored telescope. White-light polarization observations were carried out at three distinct angles using a polarizing filter placed in front of the camera system. To calibrate the intensity of the

  11. Ethan Lake

    We refine an old idea for performing fault-tolerant error correction in topological codes by simulating confining interactions between excitations. We implement confinement using an array of local classical processors that measure syndromes, broadcast messages to neighboring processors, and move excitations using received messages. The dynamics of the result

  12. Gunjun Lee, Jiwon Kim, Jaiyoung Park, Younjoo Lee

    Large Language Model (LLM) inference in production must meet stringent service-level objectives for both time-to-first-token (TTFT) and time-between-token (TBT) while maximizing throughput under fixed compute, memory, and interconnect budgets. Modern serving systems adopt stall-free scheduling techniques such as chunked prefill, which splits the processing o

  13. Moon Ye-Bin, Roy Miles, Tae-Hyun Oh, Ismail Elezi

    Image retouching not only enhances visual quality but also serves as a means of expressing personal preferences and emotions. However, existing learning-based approaches require large-scale paired data and operate as black boxes, making the retouching process opaque and limiting their adaptability to handle diverse, user- or image-specific adjustments. In th

  14. Meng Ji, Kwok-Kun Kwong

    In this paper, we uncover a novel connection between the Fenchel-Willmore inequality and a new logarithmic Sobolev inequality for mean-convex submanifolds immersed in non-negatively curved manifolds with Euclidean volume growth. Building on this connection, we establish extensions of the Fenchel-Willmore inequality to submanifolds with boundary and to comple

  15. Bheeshm Sharma, Karthikeyan Jaganathan, Balamurugan Palaniappan

    Weakly Supervised Anomaly detection (WSAD) in brain MRI scans is an important challenge useful to obtain quick and accurate detection of brain anomalies when precise pixel-level anomaly annotations are unavailable and only weak labels (e.g., slice-level) are available. In this work, we propose RASALoRE: Region Aware Spatial Attention with Location-based Rand

  16. Debashish Goswami, Kiran Maity

    We use categorical description of the invariant 2-cohomology group of Hopf algebra to compute such cohomology for two finite dimensional Hopf algebras: the group ring of $Z_8\rtimes Aut(Z_8)$ and Kac-Paljutkin algebra. For the first of these two examples, our categorical approach helps to settle the problem of computing this cohomology, which was left open i

  17. Congmin Zheng, Jiachen Zhu, Zhuoying Ou, Yuxiang Chen

    Although Large Language Models (LLMs) exhibit advanced reasoning ability, conventional alignment remains largely dominated by outcome reward models (ORMs) that judge only final answers. Process Reward Models(PRMs) address this gap by evaluating and guiding reasoning at the step or trajectory level. This survey provides a systematic overview of PRMs through t

  18. Yi-Cheng Lin, Yu-Hsuan Li Liang, Hsuan Su, Tzu-Quan Lin

    Robust ASR under domain shift is crucial because real-world systems encounter unseen accents and domains with limited labeled data. Although pseudo-labeling offers a practical workaround, it often introduces systematic, accent-specific errors that filtering fails to fix. We ask: How can we correct these recurring biases without target ground truth? We propos

  19. Qingyuan Shi, Qingwen Meng, Hao Cheng, Qing Xu

    The generation of testing and training scenarios for autonomous vehicles has drawn significant attention. While Large Language Models (LLMs) have enabled new scenario generation methods, current methods struggle to balance command adherence accuracy with the realism of real-world driving environments. To reduce scenario description complexity, these methods

  20. Artem Chernobrovkin, Marco Sälzer, François Schwarzentruber, Nicolas Troquard

    We introduce a logical language for reasoning about quantized aggregate-combine graph neural networks with global readout (ACR-GNNs). We provide a logical characterization and use it to prove that verification tasks for quantized GNNs with readout are (co)NEXPTIME-complete. This result implies that the verification of quantized GNNs is computationally intrac

  21. Shiyuan Yin, Chenjia Bai, Zihao Zhang, Junwei Jin

    Large language models (LLMs) demonstrate advanced reasoning abilities, enabling robots to understand natural language instructions and generate high-level plans with appropriate grounding. However, LLM hallucinations present a significant challenge, often leading to overconfident yet potentially misaligned or unsafe plans. While researchers have explored unc

  22. Ivan Kuznetsov, Jacopo Grassi, Dmitrii Pantiukhin, Boris Shapkin

    Large language models (LLMs) are increasingly deployed for climate-related applications, where understanding internal climatological knowledge is crucial for reliability and misinformation risk assessment. Despite growing adoption, the capacity of LLMs to recall climate normals from parametric knowledge remains largely uncharacterized. We investigate the cap

  23. Vincent Michael Sutanto, Giovanni Gatti De Giacomo, Toshiaki Nakazawa, Masaru Yamada

    This study investigates ChatGPT for Japanese-English translation, exploring simple and enhanced prompts and comparing against commercially available translation engines. Performing both automatic and MQM-based human evaluations, we found that document-level translation outperforms sentence-level translation for ChatGPT. On the other hand, we were not able to

  24. Yoshimasa Watanabe, Takahiro Oyama, Akemi Tamanai, Shaoshan Zeng

    Methanol is a seed species of complex organic molecules that is of fundamental importance in astrochemistry. Although various isotopologues of CH$_3$OH have been detected in the interstellar medium (ISM), CH$_{3}$$^{17}$OH is only tentatively detected in Sgr~B2. To confirm the presence of CH$_{3}$$^{17}$OH in the ISM and to investigate its abundance, we sear

  25. Pere Munar-Vallespir, Marc Geitz, Ángeles Vázquez-Castro, Janis Nötzel

    We study the quantum limits of the ELROI beacon concept introduced by Holmes, Weaver, and Palmer. In this concept, a satellite continuously emits a weak optical signal to broadcast its identity. Via analysis of the fundamental limits on communication introduced by Shannon, Gordon, and Holevo, we demonstrate that in such scenarios, incorporating quantum techn

  26. Alexander Herold, Daniel Sobotka, Lucian Beer, Nina Bastati

    Background: We aimed to quantify hepatic vessel volumes across chronic liver disease stages and healthy controls using deep learning-based magnetic resonance imaging (MRI) analysis, and assess correlations with biomarkers for liver (dys)function and fibrosis/portal hypertension. Methods: We assessed retrospectively healthy controls, non-advanced and advanced

  27. Alexandr Buryak, Ran J. Tessler, Mikhail Troshkin

    We give a natural definition of open Hurwitz numbers, where the weight of each ramified covering includes an integer parameter $N$ taken to the power that is equal to the number of boundary components of a Riemann surface with boundary mapping to $\mathbb{CP}^1$. We prove that the resulting sequence of partition functions, depending on $N\in\mathbb{Z}$, is a

  28. Aron Dagur Beck, Elena Vagnoni

    With increased glacial melting and the need to maintain sediment continuity for ecosystem health, sediment-laden flows through hydropower plants are becoming increasingly problematic, particularly due to erosion on runner blades and buckets. A widely used mitigation strategy is the use of filters to protect Pelton turbines. However, these filters lead to rap

  29. Ansgar Steland

    Suppose (standardized) measurements or statistics are monitored to raise an alarm when a threshold is exceeded. Often, the underlying population is heterogenous with respect to important discrete variables and thus samples may consist of imbalanced classes. We propose to use thresholds which depend on such covariates to boost the sensitivity for rare classes

  30. Xiaoshuang Ji, Zhendong Zhao, Xiaoyan Gu, Xiaojun Chen

    Parameter-efficient finetuning (PEFT) aims to mitigate the substantial computational and memory overhead involved in adapting large-scale pretrained models to diverse downstream tasks. Among numerous PEFT strategies, Low-Rank Adaptation (LoRA) has emerged as one of the most widely adopted approaches due to its robust empirical performance and low implementat

  31. Ziyang Zhu

    In this paper, we presents a method for factoring morphisms between arithmetic surfaces based on the regularity of arithmetic surfaces. Using this factorization, we derive a Riemann-Hurwitz formula satisfied by the ramification divisor and the canonical divisor on arithmetic surfaces. We also extend this formula to Arakelov theory.

  32. Dorin Bucur, Giuseppe Buttazzo, Alexis de Villeroché

    In this paper we consider the scale invariant shape functional $${\mathcal{F}}_{p,q}(\Omega)=\frac{\lambda_p^{1/p}(\Omega)}{\lambda_q^{1/q}(\Omega)},$$ where $1\le q<p\le+\infty$ and $\lambda_p(\Omega)$ (respectively $\lambda_q(\Omega)$) is the first eigenvalue of the $p$-Laplacian $-\Delta_p$ (respectively $-\Delta_q$) with Dirichlet boundary condition on $

  33. Lukas Danner, Max Hofheinz, Nicolas Bourlet, Ciprian Padurariu

    Single-photon detectors are an essential part of the toolbox of modern quantum optics for implementing quantum technologies and enabling tests of fundamental physics. The low energy of microwave photons, the natural signal path for superconducting quantum devices, makes their detection much harder than for visible light. Despite impressive progress in recent

  34. Suraj Singh Gehlot, Siddhanth Gautam, Sanhita Das

    Origami-inspired self-deployable structures offer lightweight, compact, and autonomous deployment capabilities, making them highly attractive for aerospace and defence applications, such as solar panels, antennas, and reflector systems. This paper presents finite element frameworks for simulating Miura-origami units in ABAQUS, focusing on two deployment mech

  35. Afroditi Talidou

    The FitzHugh-Nagumo equations are known to admit traveling front solutions in one spatial dimension that are nonlinearly stable. This paper concerns the stability of traveling front solutions propagating on cylindrical surfaces. It is shown that such traveling fronts are nonlinearly stable on the surface of standard cylinders of constant radius. The analysis

  36. Paul Kohl

    This work introduces the notion of unoperation $\mathfrak{Un}(\hat{O})$ of some operation $\hat{O}$. Given a valid output of $\hat{O}$, the corresponding unoperation produces a set of all valid inputs to $\hat{O}$ that produce the given output. Further, the working principle of unoperations is illustrated using the example of addition. A device providing tha

  37. Chen Huang, Wei Lu, Wenxuan Zhang

    Large Reasoning Models (LRMs) have achieved impressive performance on complex reasoning tasks by generating detailed chain-of-thought (CoT) explanations. However, these responses are often excessively long, containing redundant reasoning steps that inflate inference cost and reduce usability. Controlling the length of generated reasoning without sacrificing

  38. Károly Seller, Günter Sigl

    Astrophysical processes can contribute to magnetic fields within cosmic voids either through magnetized outflows from the astrophysical large-scale structure or through superposition of dipolar contributions from individual galaxies. Such astrophysical magnetic fields represent a foreground to possible space-filling primordial magnetic fields seeded in the e

  39. S. Mignemi

    We review an instance of noncommutative geometry based on a specific realization of the model of doubly special relativity proposed by Magueijo and Smolin (MS) on noncommutative spacetime. In particular, we discuss the Hopf algebra associated to it, which has not been considered in the literature till now. We show that the momentum sector of this model can b

  40. Akira Ito, Masanori Yamada, Daiki Chijiwa, Atsutoshi Kumagai

    Recently, Ainsworth et al. empirically demonstrated that, given two independently trained models, applying a parameter permutation that preserves the input-output behavior allows the two models to be connected by a low-loss linear path. When such a path exists, the models are said to achieve linear mode connectivity (LMC). Prior studies, including Ainsworth

  41. Kehui Liu, Zhongjie Jia, Yang Li, Zhaxizhuoma

    Data-driven robotic manipulation learning depends on large-scale, high-quality expert demonstration datasets. However, existing datasets, which primarily rely on human teleoperated robot collection, are limited in terms of scalability, trajectory smoothness, and applicability across different robotic embodiments in real-world environments. In this paper, we

  42. J. M. Alendouro Pinho, B. Amorim, Yuliy V. Bludov, J. M. Viana Parente Lopes

    The theory of open quantum systems is one of the most essential tools for the development of quantum technologies. A particular area of interest is in the optical response of solid state systems, where dissipation is introduced phenomenologically through the relaxation time approximation and its effects are usually gauged perturbatively. Analytical exact res

  43. Keren Duer-Milner, Nimrod Gavriel, Eli Galanti, Eli Tziperman

    The equatorial jets dominating the dynamics of the Jovian planets exhibit two distinct types of zonal flows: strongly eastward in the gas giants (superrotation) and strongly westward in the ice giants (subrotation). Existing theories propose different mechanisms for these patterns, but no single mechanism has successfully explained both. However, the planeta

  44. Praveen C. Srivastava, Sakshi Shukla

    In the present work, we aim to study collectivity in the Pb isotopes in the framework of nuclear shell model. We have performed shell-model calculations using KHH7B effective interaction. The model space of KHH7B interaction consists of 14 orbitals. We have reported results for even-even $^{196-206}$Pb isotopes for spectra and electromagnetic properties. The

  45. Shaohong Wang, Bin Lu, Xinyu Xiao, Hanzhi Zhong

    Collaborative visual perception methods have gained widespread attention in the autonomous driving community in recent years due to their ability to address sensor limitation problems. However, the absence of explicit depth information often makes it difficult for camera-based perception systems, e.g., 3D object detection, to generate accurate predictions. T

  46. Stanisław Pawlak, Jan Dubiński, Daniel Marczak, Bartłomiej Twardowski

    Model merging (MM) recently emerged as an effective method for combining large deep learning models. However, it poses significant security risks. Recent research shows that it is highly susceptible to backdoor attacks, which introduce a hidden trigger into a single fine-tuned model instance that allows the adversary to control the output of the final merged

  47. Zheng Xing, Junting Chen

    Radio maps are essential for enhancing wireless communications and localization. However, existing methods for constructing radio maps typically require costly calibration processes to collect location-labeled channel state information (CSI) datasets. This paper aims to recover the data collection trajectory directly from the channel propagation sequence, el

  48. Yan Meng, Kainan Chang, Yanyan Qian, Luxia Wang

    Black phosphorene (BP) has emerged as a promising platform for tunable nonlinear photonics due to its layer-dependent bandgap, high carrier mobility, and remarkable in-plane anisotropy. This study investigates the second-harmonic generation (SHG) of monolayer and bilayer BP under an external static electric field, with describing the electronic states by a t

  49. Yurang R. Kuang

    We present the discovery of a fundamental composition law governing conjugate observables in the Random Permutation Sorting System (RPSS). The law links the discrete permutation count Np and the continuous elapsed time T through a functional relation connecting the characteristic function of timing distributions to the probability generating function of perm

  50. Jinze Wang, Lu Zhang, Yiyang Cui, Tiehua Zhang

    Next point-of-interest (POI) recommendation is a key component of smart urban services, yet it remains challenging under cold-start conditions with sparse user-POI interactions. Recent LLM-based methods address this issue through either supervised fine-tuning (SFT) or in-context learning (ICL), but SFT is costly and prone to overfitting active users, while s

  51. Wei Zhang, Ding Chen, Bin Zhou

    To avoid the unpredictable phase deviations of the spaceborne phased array (SPA), this paper considers the over-the-air (OTA) phase calibration of the SPA for the low earth orbit (LEO) satellite communications, where the phase deviations of the SPA and the unknown channel are jointly estimated with multiple transmissions of the pilots. Moreover, the Cramer R

  52. Binbin Huang, Luo Luo, Yanghua Xiao, Deqing Yang

    This work proposes a novel framework based on nested evolving set processes to accelerate Personalized PageRank (PPR) computation. At each stage of the process, we employ a localized inexact proximal point iteration to solve a simplified linear system. We show that the time complexity of such localized methods is upper bounded by $\min\{\tilde{\mathcal{O}}(R

  53. Alex O. Davies, Roussel Nzoyem, Nirav Ajmeri, Telmo M. Silva Filho

    Recent research has extensively studied how large language models manipulate integers in specific arithmetic tasks, and on a more fundamental level, how they represent numeric values. These previous works have found that language model embeddings can be used to reconstruct the original values, however, they do not evaluate whether language models actually mo

  54. Haoran Ou, Kangjie Chen, Xingshuo Han, Gelei Deng

    Large Language Models (LLMs) have been augmented with web search to overcome the limitations of the static knowledge boundary by accessing up-to-date information from the open Internet. While this integration enhances model capability, it also introduces a distinct safety threat surface: the retrieval and citation process has the potential risk of exposing u

  55. Sebastian Lahs, Daniel Comparat, Fiona Kirk, Benjamin Roberts

    The existence of cosmic fields made from yet unknown light bosons is predicted in many extensions to the Standard Model. They are especially of interest as possible constituents of dark matter. To detect such light and weakly interacting fields, atomic precision measurements offer one of the most sensitive platforms. In this work, we derive which atomic obse

  56. Claudio Bonanno

    I present a large-$N$ determination of the topological susceptibility $\chi$ of $\mathrm{SU}(N)$ Yang--Mills theories using non-perturbative numerical Monte Carlo simulations of the lattice-discretized theory for $3\le N \le 6$, and adopting the Parallel Tempering on Boundary Conditions (PTBC) algorithm to bypass topological freezing for $N>3$. Thanks to thi

  57. Utku Boran Torun, Mehmet Taha Demircan, Mahmut Furkan Gön, Eray Tüzün

    Traditional bug-tracking systems rely heavily on manual reporting, reproduction, classification, and resolution, involving multiple stakeholders such as end users, customer support, developers, and testers. This division of responsibilities requires substantial coordination and human effort, widens the communication gap between non-technical users and develo

  58. Honghong Wang, Jing Deng, Rong Zheng

    This paper presents our solution to the Multimodal Personality-aware Depression Detection (MPDD) challenge at ACM MM 2025. We propose a multimodal depression detection model in the Elderly that incorporates personality characteristics. We introduce a multi-feature fusion approach based on a co-attention mechanism to effectively integrate LLDs, MFCCs, and Wav

  59. Weihuang Lin, Yiwei Ma, Jiayi Ji, Xiaoshuai Sun

    Composed Image Retrieval (CIR), which aims to find a target image from a reference image and a modification text, presents the core challenge of performing unified reasoning across visual and semantic modalities. While current approaches based on Vision-Language Models (VLMs, e.g., CLIP) and more recent Multimodal Large Language Models (MLLMs, e.g., Qwen-VL)

  60. Cheng Yang, Xuemeng Yang, Licheng Wen, Daocheng Fu

    Large Language Models have demonstrated remarkable capabilities across diverse domains, yet significant challenges persist when deploying them as AI agents for real-world long-horizon tasks. Existing LLM agents suffer from a critical limitation: they are test-time static and cannot learn from experience, lacking the ability to accumulate knowledge and contin

  61. Mohsen Ahadi, Omid Esrafilian, Florian Kaltenberger, Adeel Malik

    Channel Charting (CC) has emerged as a promising framework for data-driven radio localization, yet existing approaches often struggle to scale globally and to handle the distortions introduced by non-line-of-sight (NLoS) conditions. In this work, we propose a novel CC method that leverages Channel Impulse Response (CIR) data enriched with practical features

  62. Kevin Steijn, Vamsi Priya Goli, Enrico Antonini

    This paper presents a machine learning framework for electricity demand forecasting across diverse geographical regions using the gradient boosting algorithm XGBoost. The model integrates historical electricity demand and comprehensive weather and socioeconomic variables to predict normalized electricity demand profiles. To enable robust training and evaluat

  63. Michael Strunk

    In this paper, we are interested in the regularity of weak solutions $u\colon\Omega_T\to\mathbb{R}$ to parabolic equations of the type \begin{equation*} \partial_t u - \mathrm{div} \nabla \mathcal{F}(x,t,Du) = f\qquad\mbox{in $\Omega_T$}, \end{equation*} where $\mathcal{F}$ is only elliptic for values of $Du$ outside a bounded and convex set $E\subset \mathb

  64. Magdalena Łukowicz, Aleksandra Korzeniewska, Kamil Kalinowski, Rafał Cichowski

    The term wavefront sensor refers to the entire class of devices capable of measuring the optical wavefront of the incoming beam. Although numerous solutions have been proposed so far, recent advances in structured light have opened new development possibilities through controlled modification of optical field amplitude and phase. We present an alternative ap

  65. Qiyuan Chen, Hong Liu, Ke Ye

    We establish new lower bounds for the Tur\'an and Zarankiewicz numbers of certain apex partite hypergraphs. Given a $(d-1)$-partite $(d-1)$-uniform hypergraph $\mathcal{H}$, let $\mathcal{H}(k)$ be the $d$-partite $d$-uniform hypergraph whose $d$th part has $k$ vertices that share $\mathcal{ H}$ as a common link. We show that $ex(n,\mathcal{H}(k))=\Omega_{\m

  66. Stephen Piddock

    We unconditionally prove that it is NP-hard to compute a constant multiplicative approximation to the QUANTUM MAX-CUT problem on an unweighted graph of constant bounded degree. The proof works in two stages: first we demonstrate a generic reduction to computing the optimal value of a quantum problem, from the optimal value over product states. Then we prove

  67. Nathan Hancart

    I provide a sufficient condition under which a principal does not benefit from committing to a mechanism in economic models represented by a maximisation problem under constraints. These problems include mechanism design, principal-agent models or sender-receiver games. In principal-agent problems, this condition holds if the agent has a finite strategy spac

  68. Watcharapong Timklaypachara, Monrada Chiewhawan, Nopporn Lekuthai, Titipat Achakulvisut

    Scientific figure captions require both accuracy and stylistic consistency to convey visual information. Here, we present a domain-specific caption generation system for the 3rd SciCap Challenge that integrates figure-related textual context with author-specific writing styles using the LaMP-Cap dataset. Our approach uses a two-stage pipeline: Stage 1 combin

  69. Jian'an Zhang

    We present a white-box, risk-sensitive framework for jointly hedging SPX and VIX exposures under transaction costs and regime shifts. The approach couples an arbitrage-free market teacher with a control layer that enforces safety as constraints. On the market side, we integrate an SSVI-based implied-volatility surface and a Cboe-compliant VIX computation (in

  70. Nikita Doikov, Geovani Nunes Grapiglia

    In this work, we propose a method for minimizing non-convex functions with Lipschitz continuous $p$th-order derivatives, starting from $p \geq 1$. The method, however, only requires derivative information up to order $(p-1)$, since the $p$th-order derivatives are approximated via finite differences. To ensure oracle efficiency, instead of computing finite-di

  71. Gaurvi Goyal, Pham Cong Thuong, Arren Glover, Masayoshi Mizuno

    Human Pose Estimation is a crucial module in human-machine interaction applications and, especially since the rise in deep learning technology, robust methods are available to consumers using RGB cameras and commercial GPUs. On the other hand, event-based cameras have gained popularity in the vision research community for their low latency and low energy adv

  72. Haitao Jia, Ming He, Zimo Yin, Likang Wu

    Mobile GUI agents exhibit substantial potential to facilitate and automate the execution of user tasks on mobile phones. However, exist mobile GUI agents predominantly privilege autonomous operation and neglect the necessity of active user engagement during task execution. This omission undermines their adaptability to information dilemmas including ambiguou

  73. Rachel L. Franz, Jacob O. Wobbrock

    Today's virtual reality (VR) systems and environments assume that users have typical abilities, which can make VR inaccessible to people with physical impairments. However, there is not yet an understanding of how inaccessible locomotion techniques are, and which interactions make them inaccessible. To this end, we conducted a study in which people with and

  74. Bhargavi Srinivasan

    We use the discrete Ollivier-Ricci graph curvature with Ricci flow to examine the intrinsic geometry of financial markets through the empirical correlation graph of the NASDAQ 100 index. Our main result is the development of a technique to perform surgery on the neckpinch singularities that form during the Ricci flow of the empirical graph, using the behavio

  75. Gaofeng Li, Peisen Xu, Ruize Wang, Qi Ye

    Orientation learning plays a pivotal role in many tasks. However, the rotation group SO(3) is a Riemannian manifold. As a result, the distortion caused by non-Euclidean geometric nature introduces difficulties to the incorporation of local constraints, especially for the simultaneous incorporation of multiple local constraints. To address this issue, we prop

  76. Kazuki Egashira, Robin Staab, Thibaud Gloaguen, Mark Vero

    Model pruning, i.e., removing a subset of model weights, has become a prominent approach to reducing the memory footprint of large language models (LLMs) during inference. Notably, popular inference engines, such as vLLM, enable users to conveniently prune downloaded models before they are deployed. While the utility and efficiency of pruning methods have im

  77. Chandresh Sutariya, Nitin Singh

    The simultaneous restoration of high-frequency details and suppression of severe noise in low-light imagery presents a significant and persistent challenge in computer vision. While large-scale Transformer models like SwinIR have set the state of the art in performance, their high computational cost can be a barrier for practical applications. This paper inv

  78. Xianghong Xu, Rong Kang, Xiao He, Lei Zhang

    Cardinality estimation is a fundamental task in database systems and plays a critical role in query optimization. Despite significant advances in learning-based cardinality estimation methods, most existing approaches remain difficult to generalize to new datasets due to their strong dependence on raw data or queries, thus limiting their practicality in real

  79. Yiming Liang, Huan Yu, Zili Wang, Shuyou Zhang

    Recent advancements in AI-driven 3D model generation have leveraged cross modality, yet generating models with smooth surfaces and minimizing storage overhead remain challenges. This paper introduces a novel multi-stage framework for generating 3D models composed of parameterized primitives, guided by textual and image inputs. In the framework, A model gener

  80. Hanwen Jin, Chengcheng Xiao, Matias Herran, Emiliano Cortes

    Energetic electrons and holes generated from the decay of localized surface plasmons in metallic nanoparticles can be harnessed in nanoscale devices for photocatalysis, photovoltaics or sensing. In this work, we study the generation of such hot carriers in bimetallic Janus nanoparticles composed of Au, Ag and Cu using a recently developed atomistic modelling

  81. Hossein Safari, S. Kaveh Hedayati, Aminul Islam, Yi Yang

    Tomographic volumetric 3D printing offers layer free, rapid fabrication of objects with high design freedom, but is limited to relatively small curing volumes because of the optical constraints imposed by an assumed need for telecentricity. We present a method to virtually stitch multiple projections from different light sources to build a single workpiece.

  82. Qinglun Li, Yingqi Liu, Miao Zhang, Xiaochun Cao

    Decentralized training removes the centralized server, making it a communication-efficient approach that can significantly improve training efficiency, but it often suffers from degraded performance compared to centralized training. Multi-Gossip Steps (MGS) serve as a simple yet effective bridge between decentralized and centralized training, significantly r

  83. Wei Wang, Rong Cao, Yi Guo, Zhengyang Chen

    Flow-based generative models have greatly improved text-to-speech (TTS) synthesis quality, but inference speed remains limited by the iterative sampling process and multiple function evaluations (NFE). The recent MeanFlow model accelerates generation by modeling average velocity instead of instantaneous velocity. However, its direct application to TTS encoun

  84. Dhruv Jain, Harshit Shukla, Gautam Rajeev, Ashish Kulkarni

    Large scale Speech Language Models have enabled voice assistants capable of understanding natural spoken queries and performing complex tasks. However, existing speech benchmarks largely focus on isolated capabilities such as transcription or question answering and do not systematically evaluate agentic behavior or adversarial robustness. To address this, we

  85. Darya Baranouskaya, Andrea Cavallaro

    Object tags denote concrete entities and are central to many computer vision tasks, whereas abstract tags capture higher-level information, which is relevant for tasks that require a contextual, potentially subjective scene understanding. Object and abstract tags extracted from images also facilitate interpretability. In this paper, we explore which type of

  86. Mingyang Sun, Jiude Wei, Qichen He, Donglin Wang

    Enabling robots to perform precise and generalized manipulation in unstructured environments remains a fundamental challenge in embodied AI. While Vision-Language Models (VLMs) have demonstrated remarkable capabilities in semantic reasoning and task planning, a significant gap persists between their high-level understanding and the precise physical execution

  87. Jialu Du, Guiyang Hou, Yihui Fu, Chen Wu

    While large language models (LLMs) excel in mathematical and code reasoning, we observe they struggle with social reasoning tasks, exhibiting cognitive confusion, logical inconsistencies, and conflation between objective world states and subjective belief states. Through deteiled analysis of DeepSeek-R1's reasoning trajectories, we find that LLMs frequently

  88. Xingchen Guo Anqi Wang, Xiutong Deng, Yupeng Li, Guoan Li

    Ta2Pd3Te5 is a quasi-one-dimensional transition-metal telluride whose heavy atoms endow the material with strong spin-orbit coupling, while the Fermi level inside the bulk gap makes the low-energy electronic structure highly tunable.Theory and early experiments have already identified a wealth of emergent phases in this platform: an excitonic insulator drive

  89. Premt Cara, Kamilia Zaripova, David Bani-Harouni, Nassir Navab

    Rare genetic disease diagnosis faces critical challenges: insufficient patient data, inaccessible full genome sequencing, and the immense number of possible causative genes. These limitations cause prolonged diagnostic journeys, inappropriate treatments, and critical delays, disproportionately affecting patients in resource-limited settings where diagnostic

  90. Pengkun Jiao, Yiming Jin, Jianhui Yang, Chenhe Dong

    Query-product relevance prediction is vital for AI-driven e-commerce, yet current LLM-based approaches face a dilemma: SFT and DPO struggle with long-tail generalization due to coarse supervision, while traditional RLVR suffers from sparse feedback that fails to correct intermediate reasoning errors. We propose Stepwise Hybrid Examination (SHE), an RL framew

  91. Oskar Bohn Lassen, Serio Angelo Maria Agriesti, Filipe Rodrigues, Francisco Camara Pereira

    Climate policy studies require models that capture the combined effects of multiple greenhouse gases on global temperature, but these models are computationally expensive and difficult to embed in reinforcement learning. We present a multi-agent reinforcement learning (MARL) framework that integrates a high-fidelity, highly efficient climate surrogate direct

  92. Ding-Ming Huang, Jian-Huan Wang, Jie-Yin Zhang, Yuan Yao

    An atomically flat interface is achieved between face-centered cubic Al and diamond lattice Ge via molecular beam epitaxy (MBE). Based on the measurements of scanning tunneling microscopy (STM), we demonstrate an atomically resolved lateral periodic change of the electron reflectivity at the Al/Ge interface. The variation of electron reflectivity is up to 24

  93. Taiki Shibata, Kenichi Shimizu

    For coalgebras $C$ and $D$, Takeuchi proved that the category of linear functors from $\mathfrak{M}^C$ to $\mathfrak{M}^D$ preserving small coproducts is equivalent to the category of $C$-$D$-bicomodules, where $\mathfrak{M}^C$ for a coalgebra $C$ means the category of right $C$-comodules. We formulate and prove an equivariant version of this result for modu

  94. Xiangtao Meng, Tianshuo Cong, Li Wang, Wenyu Chen

    Large Language Models (LLMs) are increasingly deployed in high-stakes settings, where they face diverse risks. Numerous defense strategies have been proposed to mitigate these risks, but they are almost always evaluated in isolation. This isolated view leaves a critical question open: does mitigating one risk inadvertently change a model&#39;s exposure to ot

  95. Yaning Li, Ke Zhao, Shucheng Zheng, Xingyu Chen

    Lost architectural heritage presents interpretive challenges due to vanished structures and fragmented historical records. Using Hanyuan Hall of the Tang dynasty's Daming Palace as a case study, we conducted a formative investigation with archaeologists, heritage administrators, and visitors to identify key issues in current interpretation practices. We foun

  96. Kelley M. Hess, John Hibbard, Jennifer Donovan Meyer, Hansung B. Gim

    We present ALMA CO observations of 14 HI-detected galaxies from the CHILES survey found in a cosmic over-density at z~0.12. This is the largest collection of spatially resolved CO + HI observations beyond the local Universe (z>0.05) to date. While the HI-detected parent sample spans a range of stellar masses, star formation rates (SFR), and environments, we

  97. A. Pathania, K. K. Singh, S. K. Singh, A. Tolamatti

    The Large Area Telescope (LAT) on board the \emph{Fermi} Gamma-ray Space Telescope has been continuously providing good quality survey data of the entire sky in the high energy range from 30 MeV to 500 GeV and above since August 2008. A succession of gamma-ray source catalogs is published after a comprehensive analysis of the \emph{Fermi}--LAT data. The most

  98. Seungsu Han, Juyoung Hwang, Won Chang

    Normalizing flows with a Gaussian base provide a computationally efficient way to approximate posterior distributions in Bayesian inference, but they often struggle to capture complex posteriors with multimodality and heavy tails. We propose a stick-breaking mixture base with component-wise tail adaptation (StiCTAF) for posterior approximation. The method fi

  99. Jiabei Cheng, Changxi Chi, Jingbo Zhou, Hongyi Xin

    In single-cell perturbation prediction, a central task is to forecast the effects of perturbing a gene unseen in the training data. The efficacy of such predictions depends on two factors: (1) the similarity of the target gene to those covered in the training data, which informs model (epistemic) uncertainty, and (2) the quality of the corresponding training

  100. Nhu Ngoc Hoang, Ngoc Hoa Pham, Viet Phuong Hoang, Esteban Zimányi

    The analytics of spatiotemporal data is increasingly important for mobility analytics. Despite extensive research on moving object databases (MODs), few systems are ready on production or lightweight enough for analytics. MobilityDB is a notable system that extends PostgreSQL with spatiotemporal data, but it inherits complexity of the architecture as well. I