Skip to content

March 2025 arXiv papers — page 132

Showing 13,10113,200 of 23,633 papers

  1. Zichen Tang, Yuan Yao, Miaomiao Cui, Liefeng Bo

    Text-guided 3D human generation has advanced with the development of efficient 3D representations and 2D-lifting methods like Score Distillation Sampling (SDS). However, current methods suffer from prolonged training times and often produce results that lack fine facial and garment details. In this paper, we propose GaussianIP, an effective two-stage framewo

  2. Indukuru Ramesh Reddy, M. Kaltak, Bongjae Kim

    In this study, we present a systematic comparison of various approaches within the constrained random-phase approximation (cRPA) for calculating the Coulomb interaction parameter $U$. While defining the correlated space is straightforward for disentangled bands, the situation is more complex for entangled bands, where different projection schemes from hybrid

  3. Bei Chen, Xiaowen Xiong, Renheng Zhang, Yitang Dai

    Photonic neural networks have been considered as the promising candidates for next-generation neuromorphic computation, aiming to break both the power consumption wall and processing speed boundary of state-to-date digital computing architectures. Optics has shown its advantages in parallelism and linear manipulation. However, the lack of low-power and high-

  4. Lexin Fang, Yunyang Xu, Xiang Ma, Xuemei Li

    Deep learning has achieved significant advancements in medical image segmentation, but existing models still face challenges in accurately segmenting lesion regions. The main reason is that some lesion regions in medical images have unclear boundaries, irregular shapes, and small tissue density differences, leading to label ambiguity. However, the existing m

  5. Yu Voon Ng, Ting-Wen Lan, J. Xavier Prochaska, Amélie Saintonge

    We investigate the relationships between the cool circumgalactic medium (CGM), traced by Ca II absorption lines, and galaxy properties at $z<0.4$ using $\sim900{,}000$ galaxy-quasar pairs within $200\,\rm kpc$ from the Year 1 data of the Dark Energy Spectroscopic Instrument (DESI). This large data set enables us to obtain composite spectra with sensitivity r

  6. Yao Xu, Grace Nansamba, Anthony Skjellum, Gene Cooperman

    There is new momentum behind an interoperable ABI for MPI, which will be a major component of MPI-5. This capability brings true separation of concerns to a running MPI computation. The linking and compilation of an MPI application becomes completely independent of the choice of MPI library. The MPI application is compiled once, and runs everywhere. This ABI

  7. Mikhail Shkolnikov

    Tropical caustic of a convex domain on the plane is a canonically associated tropical analytic curve inside the domain. In this note we give a graphical proof for the classification of its intermediate vertices, implying in particular that they are always trivalent. Apart from that we explain how various known examples of tropical caustics are constructed an

  8. Yutaka Hosotani, Shuichiro Funatsu, Hisaki Hatanaka, Yuta Orikasa

    In the $SO(5) \times U(1) \times SU(3)$ gauge-Higgs unification in the Randall-Sundrum warped space the mixing in the $W$ couplings in the quark sector is induced by masses of $SO(5)$ singlet fermions which are responsible for splitting masses of down-type quarks from those of up-type quarks in each generation. We show that the observed Cabibbo-Kobayashi-Mas

  9. Hyeon Woo Park, Shu Zhang, Peter Meisenheimer, Maya Ramesh

    Multiferroic materials, characterized by the occurrence of two or more ferroic properties, hold potential in future technological applications and also exhibit intriguing phenomena caused by the interplay of multiple orders. One such example is the formation of spin cycloid structures within multiferroic materials, which we investigate in this work by focusi

  10. Manato Sakai, Yasuhiro Yamaguchi

    Recently, a number of exotic hadrons have been reported in the experiments, and most of these states lie slightly below the threshold. Therefore, these states are considered to be hadronic molecules composed of mesons or baryons. In 2022, the doubly charmed tetraquark $T_{cc}$ was reported by the LHCb experiment, which is considered to be composed of two hea

  11. Hao Liu, Pengyu Guo, Siyuan Yang, Zeqing Jiang

    With the continuous advancement of human exploration into deep space, intelligent perception and high-precision segmentation technology for on-orbit multi-spacecraft targets have become critical factors for ensuring the success of modern space missions. However, the complex deep space environment, diverse imaging conditions, and high variability in spacecraf

  12. Guihong Li, Mehdi Rezagholizadeh, Mingyu Yang, Vikram Appia

    Multi-head latent attention (MLA) is designed to optimize KV cache memory through low-rank key-value joint compression. Rather than caching keys and values separately, MLA stores their compressed latent representations, reducing memory overhead while maintaining the performance. While MLA improves memory efficiency without compromising language model accurac

  13. Vijay Bhattiprolu, Venkatesan Guruswami, Xuandi Ren

    We give simple deterministic reductions demonstrating the NP-hardness of approximating the nearest codeword problem and minimum distance problem within arbitrary constant factors (and almost-polynomial factors assuming NP cannot be solved in quasipolynomial time). The starting point is a simple NP-hardness result without a gap, and is thus "PCP-free." Our ap

  14. Ruojing Zhao, Yifei Xu, Songjie Yang, Hua Chen

    In the development of wireless communication technology, multiple-input multiple-output (MIMO) technology has emerged as a key enabler, significantly enhancing the capacity of communication systems. However, traditional MIMO systems, which rely on fixed-position antennas (FPAs) with spacing limitations, cannot fully exploit the channel variations in the cont

  15. Yijia Xu, Jianzhong Ju, Jian Luan, Jinshi Cui

    The raster-ordered image token sequence exhibits a significant Euclidean distance between index-adjacent tokens at line breaks, making it unsuitable for autoregressive generation. To address this issue, this paper proposes Direction-Aware Diagonal Autoregressive Image Generation (DAR) method, which generates image tokens following a diagonal scanning order.

  16. Ben O'Neill

    We examine the Gaussian hypergeometric beta distribution and look at the effect of having an additional term in the density kernel relative to the standard beta distribution. We reparameterise and classify this distribution into left and right directional variants using parameters that give a simple and symmetrical representation of the directional push/pull

  17. Matthew Khoriaty, Andrii Shportko, Gustavo Mercier, Zach Wood-Doughty

    Recent developments in Large Language Model (LLM) capabilities have brought great potential but also posed new risks. For example, LLMs with knowledge of bioweapons, advanced chemistry, or cyberattacks could cause violence if placed in the wrong hands or during malfunctions. Because of their nature as near-black boxes, intuitive interpretation of LLM interna

  18. Jie Liu, Yiwei Zhang, Yuan Sheng, Yujia Lou

    This study proposes a dynamic rule data mining algorithm based on an improved Transformer architecture, aiming to improve the accuracy and efficiency of rule mining in a dynamic data environment. With the increase in data volume and complexity, traditional data mining methods are difficult to cope with dynamic data with strong temporal and variable character

  19. Yongyi Jia, Shu Miao, Jiayu Wu, Ming Yang

    While magnetic micro-robots have demonstrated significant potential across various applications, including drug delivery and microsurgery, the open issue of precise navigation and control in complex fluid environments is crucial for in vivo implementation. This paper introduces a novel flow-aware navigation and control strategy for magnetic micro-robots that

  20. Songjie Yang, Jiahe Guo, Zilin He, Boyu Ning

    Flexible-geometry arrays have garnered much attention in wireless communications, which dynamically adjust wireless channels to improve the system performance. In this paper, we propose a novel flexible-geometry array for a $360^\circ$ coverage, named flxible cylindrical array (FCLA), comprised of multiple flexible circular arrays (FCAs). The elements in eac

  21. Hongbin Lin, Zilu Guo, Yifan Zhang, Shuaicheng Niu

    In autonomous driving, vision-centric 3D detection aims to identify 3D objects from images. However, high data collection costs and diverse real-world scenarios limit the scale of training data. Once distribution shifts occur between training and test data, existing methods often suffer from performance degradation, known as Out-of-Distribution (OOD) problem

  22. Mika Juvela

    Radiative transfer effects need to be taken into account when analysing spectral line observations. When the data are not sufficient for detailed modelling, simpler methods are needed. The escape probability formalism (EPF) is one such tool. We wish to quantify the model errors in the EPF analysis of interstellar clouds and cores. We introduce PEP, a paralle

  23. Gökhan Özbulak, Oscar Jimenez-del-Toro, Maíra Fatoretto, Lilian Berton

    The evaluation of fairness models in Machine Learning involves complex challenges, such as defining appropriate metrics, balancing trade-offs between utility and fairness, and there are still gaps in this stage. This work presents a novel multi-objective evaluation framework that enables the analysis of utility-fairness trade-offs in Machine Learning systems

  24. Weifeng Shang, Jose Abel Castellanos Joo, Chenqi Mou, Deepak Kapur

    New results on computing certificates of strictly positive polynomials in Archimedean quadratic modules are presented. The results build upon (i) Averkov's method for generating a strictly positive polynomial for which a membership certificate can be more easily computed than the input polynomial whose certificate is being sought, and (ii) Lasserre's method

  25. Kristin Qi, Youxiang Zhu, Xiaohui Liang

    We present our approach to the PerAnsSumm Shared Task, which involves perspective span identification and perspective-aware summarization in community question-answering (CQA) threads. For span identification, we adopt ensemble learning that integrates three transformer models through averaging to exploit individual model strengths, achieving an 82.91% F1-sc

  26. Kaixuan Jiang, Yang Liu, Weixing Chen, Jingzhou Luo

    Embodied Question Answering (EQA) is a challenging task in embodied intelligence that requires agents to dynamically explore 3D environments, actively gather visual information, and perform multi-step reasoning to answer questions. However, current EQA approaches suffer from critical limitations in exploration efficiency, dataset design, and evaluation metri

  27. Hanbyul Song, Miguel F. Santos Silva, Jaume Suau, Luis Espinosa-Anke

    Understanding why people trust or distrust one another, institutions, or information is a complex task that has led scholars from various fields of study to employ diverse epistemological and methodological approaches. Despite the challenges, it is generally agreed that the antecedents of trust (and distrust) encompass a multitude of emotional and cognitive

  28. Jun Yu, Yunxiang Zhang, Xilong Lu, Yang Zheng

    In this report, we present our solution for the Action Unit (AU) Detection Challenge, in 8th Competition on Affective Behavior Analysis in-the-wild. In order to achieve robust and accurate classification of facial action unit in the wild environment, we introduce an innovative method that leverages audio-visual multimodal data. Our method employs ConvNeXt as

  29. Vishal Gandhi, Sagar Gandhi

    The rise of large language models (LLMs) has revolutionized natural language processing (NLP), yet the influence of prompt sentiment, a latent affective characteristic of input text, remains underexplored. This study systematically examines how sentiment variations in prompts affect LLM-generated outputs in terms of coherence, factuality, and bias. Leveragin

  30. Guillermo Nuñez Ponasso

    We study the maximum absolute value of the determinant of matrices with entries in the set of $\ell$-th roots of unity; this is a generalization of $D$-optimal designs and Hadamard's maximal determinant problem, which involves $\pm 1$ matrices. For general values of $\ell$, we give sharpened determinantal upper bounds and constructions of matrices of large d

  31. Yanwei Huang, Wesley Hanwen Deng, Sijia Xiao, Motahhare Eslami

    Generative text-to-image (T2I) models are known for their risks related such as bias, offense, and misinformation. Current AI auditing methods face challenges in scalability and thoroughness, and it is even more challenging to enable auditors to explore the auditing space in a structural and effective way. Vipera employs multiple visual cues including a scen

  32. Songjie Yang, Zihang Wan, Boyu Ning, Weidong Mei

    Typical reconfigurable intelligent surface (RIS) implementations include metasurfaces with almost passive unit elements capable of reflecting their incident waves in controllable ways, enhancing wireless communications in a cost-effective manner. In this paper, we advance the concept of intelligent metasurfaces by introducing a flexible array geometry, terme

  33. Chen Zhong, Xufeng Zhou, Lan Tang, Mengting Lou

    In this paper, we investigate a distributed multi-input multi-output and orthogonal frequency division multiplexing (MIMO-OFDM) dual-function radar-communication (DFRC) system, which enables simultaneous communication and sensing in different subcarrier sets. To obtain the best tradeoff between communication and sensing performance, we first derive Cramer-Ra

  34. Zhenlei Wanga, Yaxi Yua, Mengkai Qin, Hao Jiang

    Bubble formation in electrochemical system often hinders reaction efficiency by reducing active surface area and obstructing mass transfer, yet the mechanisms governing their nanoscale nucleation dynamics and impact remains unclear. In this study, we used molecular dynamics simulations to explore nanobubble nucleation and reaction rates during water electrol

  35. Julia Gersey, Rose Allegrette, Joshua Lian, Zawad Munshi

    The growing homelessness crisis in the U.S. presents complex social, economic, and public health challenges, straining shelters, healthcare, and social services while limiting effective interventions. Traditional assessment methods struggle to capture its dynamic, dispersed nature, highlighting the need for scalable, data-driven detection. This survey explor

  36. Qinhua Guo, Lizhou Yang, Yawen Gan, Jingyang Zhang

    Micro-transfer printing is an assembly technology that enables large-scale integration of diverse materials and components from micro- to nano-scale. However, traditional micro-transfer printing technologies lack dynamic selectivity, limiting capabilities in sorting and repairing materials and components for effective yield management during large-scale manu

  37. Yifang Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi

    The key-value (KV) cache in the tensor version of transformers presents a significant bottleneck during inference. While previous work analyzes the fundamental space complexity barriers in standard attention mechanisms [Haris and Onak, 2025], our work generalizes the space complexity barriers result to tensor attention version. Our theoretical contributions

  38. Song Cao, Taikun Zhu, Kai Jin

    This paper addresses resource allocation problem with a separable objective function under a single linear constraint, formulated as maximizing $\sum_{j=1}^{n}R_j(x_j)$ subject to $\sum_{j=1}^{n}x_j=k$ and $x_j\in\{0,\dots,m\}$. While classical dynamic programming approach solves this problem in $O(n^2m^2)$ time, we propose a regret-enabled greedy algorithm

  39. Jiawei Wang, Xuedong Hu, Herbert F Fotso

    We study the low-energy spectrum of a single hole confined in a planar Ge quantum dot (QD) within the effective-mass formalism. The QD is sandwiched between two GeSi barriers of finite potential height grown along the [001] direction. To treat this finite barrier problem, we adopt an independent-band approach in dealing with boundary conditions. The effects

  40. Shivam Dubey, Mrinal Kanti Roychowdhury, Saurabh Verma

    For a given $r\in (0, +\infty)$, the quantization dimension of order $r$, if it exists, denoted by $D_r(μ)$, of a Borel probability measure $μ$ on ${\mathbb R}^d$ represents the speed how fast the $n$th quantization error of order $r$ approaches to zero as the number of elements $n$ in an optimal set of $n$-means for $μ$ tends to infinity. If $D_r(μ)$ does n

  41. Lei Qin, Ye Pu

    Optimization problems involving the minimization of a finite sum of smooth, possibly non-convex functions arise in numerous applications. To achieve a consensus solution over a network, distributed optimization algorithms, such as \textbf{EXTRA} (decentralized exact first-order algorithm), have been proposed to address these challenges. In this paper, we ana

  42. Avinash Madasu, Vasudev Lal, Phillip Howard

    CLIP is one of the most popular foundation models and is heavily used for many vision-language tasks, yet little is known about its inner workings. As CLIP is increasingly deployed in real-world applications, it is becoming even more critical to understand its limitations and embedded social biases to mitigate potentially harmful downstream consequences. How

  43. Xiaoqi Zhang, Zhitong Ni, Weijie Yuan, J. Andrew Zhang

    Orthogonal Time Frequency Space (OTFS) modulation has recently attracted significant interest due to its potential for enabling reliable communication in high-mobility environments. However, the effectiveness of OTFS receivers relies on the inherent characteristic of the Delay-Doppler (DD) domain channel, where the sparsity of the discretized channel varies

  44. Asifullah Khan, Laiba Asmatullah, Anza Malik, Shahzaib Khan

    Self-supervised learning is a machine learning approach that generates implicit labels by learning underlined patterns and extracting discriminative features from unlabeled data without manual labelling. Contrastive learning introduces the concept of "positive" and "negative" samples, where positive pairs (e.g., variation of the same image/object) are brough

  45. Bo Li, Feifei Zhang, Yu Pang, Jinyu Hu

    The efficiency of silicon solar cells gradually decreases in various environments, with humidity being a key factor contributing to this decline through moisture-induced degradation (MID) involving multiple mechanisms including encapsulant hydrolysis and metal ion migration. Among these mechanisms, the role of water-derived hydrogen and oxygen interstitial d

  46. Arnab Bhattacharyya, Weiming Feng, Piyush Srivastava

    The total variation distance is a metric of central importance in statistics and probability theory. However, somewhat surprisingly, questions about computing it algorithmically appear not to have been systematically studied until very recently. In this paper, we contribute to this line of work by studying this question in the important special case of multi

  47. Zeliang Wu, Jinxian Guo, Zhifei Yu, Wenfeng Huang

    High-dimensional broadband quantum memory significantly expands quantum information processing capabilities, but the memory efficiency becomes insufficient when extended to high dimensions. We demonstrate an efficient quantum memory for hyper-dimensional photons encoded with orbital angular momentum (OAM) and spin angular momentum (SAM). OAM information is e

  48. Wenbang Deng, Xieyuanli Chen, Qinghua Yu, Yunze He

    Semantic segmentation is a key technique that enables mobile robots to understand and navigate surrounding environments autonomously. However, most existing works focus on segmenting known objects, overlooking the identification of unknown classes, which is common in real-world applications. In this paper, we propose a feature-oriented framework for open-set

  49. He Zhang, Xinyi Fu, John M. Carroll

    Traditional image annotation tasks rely heavily on human effort for object selection and label assignment, making the process time-consuming and prone to decreased efficiency as annotators experience fatigue after extensive work. This paper introduces a novel framework that leverages the visual understanding capabilities of large multimodal models (LMMs), pa

  50. Mikihiro Fujii

    We consider the stationary problem for the quasi-geostrophic equation with the critical and super-critical dissipation and prove the unique existence of small solutions for given small external force in the scaling critical Sobolev spaces framework. Moreover, we also show that the data-to-solution map is continuous. Since the critical and super-critical case

  51. Weichen Zhang, Zile Zhou, Xin Zeng, Xuchen Liu

    Spatial reasoning is a fundamental capability of multimodal large language models (MLLMs), yet their performance in open aerial environments remains underexplored. In this work, we present Open3D-VQA, a novel benchmark for evaluating MLLMs' ability to reason about complex spatial relationships from an aerial perspective. The benchmark comprises 73k QA pairs

  52. Yuan Liu, Saihui Hou, Saijie Hou, Jiabao Du

    Image Difference Captioning (IDC) aims to generate natural language descriptions of subtle differences between image pairs, requiring both precise visual change localization and coherent semantic expression. Despite recent advancements, existing datasets often lack breadth and depth, limiting their applicability in complex and dynamic environments: (1) from

  53. Mikihiro Fujii, Tsukasa Iwabuchi

    We consider the stationary problem for the quasi-geostrophic equation on the whole plane and investigate its well-posedness and ill-posedness. In[Fujii, Ann. PDE 10, 10 (2024)], it was shown that the two-dimensional stationary Navier--Stokes equations are ill-posed in the critical Besov spaces $\dot B_{p,1}^{\frac{2}{p}-1}(\mathbb{R}^2)$ with $1 \leq p \leq

  54. Ganlong Zhao, Guanbin Li, Jia Pan, Yizhou Yu

    Aerial Vision-and-Language Navigation (Aerial VLN) aims to obtain an unmanned aerial vehicle agent to navigate aerial 3D environments following human instruction. Compared to ground-based VLN, aerial VLN requires the agent to decide the next action in both horizontal and vertical directions based on the first-person view observations. Previous methods strugg

  55. Scott C Evans, Nathan Dahlin, Ibrahima Ndiaye, Sachini Piyoni Ekanayake

    We propose a disruptive paradigm to actively place and schedule TWhrs of parallel AI jobs strategically on the grid, at distributed, grid-aware high performance compute data centers (HPC) capable of using their massive power and energy load to stabilize the grid while reducing grid build-out requirements, maximizing use of renewable energy, and reducing Gree

  56. Ziyu Huang, Chuanfei Dong, Liang Wang

    Nonlinear plasma physics problems are usually simulated through comprehensive modeling of phase space. The extreme computational cost of such simulations has motivated the development of multi-moment fluid models. However, a major challenge has been finding a suitable fluid closure for these fluid models. Recent developments in physics-informed machine learn

  57. Yi Zhang, Qiang Zhang, Xiaozhu Ju, Zhaoyang Liu

    While multimodal large language models (MLLMs) have made groundbreaking progress in embodied intelligence, they still face significant challenges in spatial reasoning for complex long-horizon tasks. To address this gap, we propose EmbodiedVSR (Embodied Visual Spatial Reasoning), a novel framework that integrates dynamic scene graph-guided Chain-of-Thought (C

  58. Yifan Liu, Xun Xu, Shijie Li, Jingyi Liao

    Multi-camera systems provide richer contextual information for industrial anomaly detection. However, traditional methods process each view independently, disregarding the complementary information across viewpoints. Existing multi-view anomaly detection approaches typically employ data-driven cross-view attention for feature fusion but fail to leverage the

  59. Hsiang Lee, Shinichi Nishihaya, Markus Kriener, Jun Fujioka

    Recent observation of the in-plane anomalous Hall effect in magnetic Weyl semimetal EuCd2Sb2 has drawn attention to out-of-plane orbital magnetization induced by an in-plane field component. Here we study EuZn2Sb2, a sister compound of EuCd2Sb2, to demonstrate sensitive changes of the in-plane anomalous Hall effect on the band modulation. The Hall resistivit

  60. Haihong Zhao, Zhixun Li, Chenyi Zi, Aochuan Chen

    Graph learning plays a vital role in mining and analyzing complex relationships within graph data and has been widely applied to real-world scenarios such as social, citation, and e-commerce networks. Foundation models in computer vision (CV) and natural language processing (NLP) have demonstrated remarkable cross-domain capabilities that are equally signifi

  61. Sixiang Ye, Zeyu Sun, Guoqing Wang, Liwei Guo

    Code generation has emerged as a key task to automate software development by converting high-level descriptions into executable code. Large language models (LLMs) excel at this but depend heavily on input prompt quality.Manual prompt engineering can be time-consuming and inconsistent, limiting LLM effectiveness. This paper introduces Prochemy, an innovative

  62. Zhou Fang, Hanlu Zhang, Jacky He, Zhen Qi

    This study aims to develop an efficient and accurate model for detecting malicious comments, addressing the increasingly severe issue of false and harmful content on social media platforms. We propose a deep learning model that combines BERT and BiLSTM. The BERT model, through pre-training, captures deep semantic features of text, while the BiLSTM network ex

  63. Yangyang Xie, Cheng Hu, Nicolas Baumann, Edoardo Ghignone

    Autonomous drifting is a complex challenge due to the highly nonlinear dynamics and the need for precise real-time control, especially in uncertain environments. To address these limitations, this paper presents a hierarchical control framework for autonomous vehicles drifting along general paths, primarily focusing on addressing model inaccuracies and mitig

  64. Liwei Guo, Sixiang Ye, Zeyu Sun, Xiang Chen

    Large Language Models (LLMs) have demonstrated remarkable performance in code completion. However, the training data used to develop these models often contain a significant amount of buggy code. Yet, it remains unclear to what extent these buggy instances influence LLMs' performance when tackling bug-prone code completion tasks. To fill this gap, this paper

  65. Pingrui Zhang, Xianqiang Gao, Yuhan Wu, Kehui Liu

    In mobile manipulation, navigation and manipulation are often treated as separate problems, resulting in a significant gap between merely approaching an object and engaging with it effectively. Many navigation approaches primarily define success by proximity to the target, often overlooking the necessity for optimal positioning that facilitates subsequent ma

  66. Wuwei Huang, Renren Jin, Wen Zhang, Jian Luan

    Recent studies on end-to-end speech translation(ST) have facilitated the exploration of multilingual end-to-end ST and end-to-end simultaneous ST. In this paper, we investigate end-to-end simultaneous speech translation in a one-to-many multilingual setting which is closer to applications in real scenarios. We explore a separate decoder architecture and a un

  67. Zhiyu Dong, Patrick A. Lee

    Can strong repulsive interactions be shown to give rise to pairing in a controlled way? We find that for a single flavor polarized band, there is a small expansion parameter in the low density limit, once the Bloch wavefunction form factor is taken into account. A perturbative expansion is possible, even if the interaction is much stronger than the Fermi ene

  68. Taehwan Lee, Kyeongkook Seo, Jaejun Yoo, Sung Whan Yoon

    Flat minima, known to enhance generalization and robustness in supervised learning, remain largely unexplored in generative models. In this work, we systematically investigate the role of loss surface flatness in generative models, both theoretically and empirically, with a particular focus on diffusion models. We establish a theoretical claim that flatter m

  69. Joseph Oglio, Mikhail Nesterenko, Gokarna Sharma

    We present SmartShards: a new sharding algorithm for improving Byzantine tolerance and churn resistance in blockchains. Our algorithm places a peer in multiple shards to create an overlap. This simplifies cross-shard communication and shard membership management. We describe SmartShards, prove it correct and evaluate its performance. We propose several Smart

  70. Pinghui Huang, Xue-Ning Bai

    Dust concentration in protoplanetary disks (PPDs) is the first step towards planetesimal formation, a crucial yet highly uncertain stage in planet formation. Although the streaming instability (SI) is widely recognized as a powerful mechanism for planetesimal formation, its properties can be sensitive to the gas dynamical environment. The outer region of PPD

  71. Zixiao Ma, Baosen Zhang

    The growing integration of inverter-based resources (IBRs) into modern power systems poses significant challenges for maintaining reliable operation under dynamic and constrained conditions. This paper focuses on the power tracking problem for grid-connected IBRs, addressing the complexities introduced by voltage and power factor constraints. Voltage constra

  72. Xueyang Zhou, Guiyao Tie, Guowen Zhang, Weidong Wang

    The rise of Large Reasoning Models (LRMs) signifies a paradigm shift toward advanced computational reasoning. Yet, this progress disrupts traditional agent frameworks, traditionally anchored by execution-oriented Large Language Models (LLMs). To explore this transformation, we propose the LaRMA framework, encompassing nine tasks across Tool Usage, Plan Desig

  73. Hongyang Wei, Shuaizheng Liu, Chun Yuan, Lei Zhang

    By leveraging the generative priors from pre-trained text-to-image diffusion models, significant progress has been made in real-world image super-resolution (Real-ISR). However, these methods tend to generate inaccurate and unnatural reconstructions in complex and/or heavily degraded scenes, primarily due to their limited perception and understanding capabil

  74. Haotian Tan, Yuan-Hua Ni

    Time-optimal trajectory planning and control is central for autonomous vehicles, yet its application and real-time deployment confronts two fundamental challenges: the non-convexity of optimal control problems and the unpredictable computation time inherent to nonlinear programming. To address these challenges, we propose a hierarchical convex optimization f

  75. Zhenguang Liu, Chao Shuai, Shaojing Fan, Ziping Dong

    Diffusion models have achieved remarkable success in novel view synthesis, but their reliance on large, diverse, and often untraceable Web datasets has raised pressing concerns about image copyright protection. Current methods fall short in reliably identifying unauthorized image use, as they struggle to generalize across varied generation tasks and fail whe

  76. Kelu Yao, Nuo Xu, Rong Yang, Yingying Xu

    This paper introduces a holistic vision-language foundation model tailored for remote sensing, named Falcon. Falcon offers a unified, prompt-based paradigm that effectively executes comprehensive and complex remote sensing tasks. Falcon demonstrates powerful understanding and reasoning abilities at the image, region, and pixel levels. Specifically, given sim

  77. Hyunwoo Park, Baekryun Seong, Sang-Ki Ko

    In cooperative multi-agent reinforcement learning (MARL), the permutation problem where the state space grows exponentially with the number of agents reduces sample efficiency. Additionally, many existing architectures struggle with scalability, relying on a fixed structure tied to a specific number of agents, limiting their applicability to environments wit

  78. Chaoyun Zhang, Shilin He, Liqun Li, Si Qin

    Large language models (LLMs) have evolved beyond simple text generation to power software agents that directly translate natural language commands into tangible actions. While API-based LLM agents initially rose to prominence for their robust automation capabilities and seamless integration with programmatic endpoints, recent progress in multimodal LLM resea

  79. Leqi Lin, Xingyu Zhou, Kaiyuan Yang, Xizhong Chen

    Pharmaceutical process design and development for generic, innovative, or personalized drugs have always been a time-consuming, costly, rigorous process, that involves multi-stage evaluation for better quality control and assurance. Large language models (LLMs), a type of generative artificial intelligence system, can augment laboratory research in the pharm

  80. Bin Liu, Xiaohong Liu, Qin Luo, Ziqiao Shang

    Pairwise learning underpins implicit collaborative filtering, yet its effectiveness is often hindered by sparse supervision, noisy interactions, and popularity-driven exposure bias. In this paper, we propose Variational Bayesian Personalized Ranking (VarBPR), a tractable variational framework for implicit-feedback pairwise learning that offers principled exp

  81. I. Bentley, J. Tedder, M. Gebran, A. Paul

    This paper describes the development of the Four Model Tree Ensemble (FMTE). The FMTE is a composite of machine learning models trained on experimental binding energies from the Atomic Mass Evaluation (AME) 2012. The FMTE predicts binding energy values for all nuclei with N > 7 and Z > 7 from AME 2020 with a standard deviation of 76 keV and a mean average de

  82. Peter Böhm, Pauline Pounds, Archie C. Chapman

    Deep reinforcement learning (DRL) has had success in virtual and simulated domains, but due to key differences between simulated and real-world environments, DRL-trained policies have had limited success in real-world applications. To assist researchers to bridge the \textit{sim-to-real gap}, in this paper, we describe a low-cost physical inverted pendulum a

  83. Ziqi Wang, Derek Hua, Wenjun Jiang, Tianwei Xing

    Respiration waveforms are increasingly recognized as important biomarkers, offering insights beyond simple respiration rates, such as detecting breathing irregularities for disease diagnosis or monitoring breath patterns to guide rehabilitation training. Previous works in wireless respiration monitoring have primarily focused on estimating respiration rate,

  84. Tatia Kiliptari, Dmitrii L. Maslov

    The $T^2$-scaling of resistivity with temperature is often viewed as a classic hallmark of a Fermi-liquid (FL) behavior in metals. However, if umklapp scattering is suppressed, this scaling is not universally guaranteed to occur. In this case, the resistivity behavior is influenced by several factors, such as dimensionality (two vs. three), topology (simply-

  85. Wenhao Jiang, Duo Li, Menghan Hu, Chao Ma

    In the field of autonomous driving, end-to-end deep learning models show great potential by learning driving decisions directly from sensor data. However, training these models requires large amounts of labeled data, which is time-consuming and expensive. Considering that the real-world driving data exhibits a long-tailed distribution where simple scenarios

  86. Jordan S. Ellenberg, Cristofero S. Fraser-Taliente, Thomas R. Harvey, Karan Srivastava

    We present a new implementation of the LLM-driven genetic algorithm {\it funsearch}, whose aim is to generate examples of interest to mathematicians and which has already had some success in problems in extremal combinatorics. Our implementation is designed to be useful in practice for working mathematicians; it does not require expertise in machine learning

  87. Heng Wang, Yotaro Shimose, Shingo Takamatsu

    Advertising banners are critical for capturing user attention and enhancing advertising campaign effectiveness. Creating aesthetically pleasing banner designs while conveying the campaign messages is challenging due to the large search space involving multiple design elements. Additionally, advertisers need multiple sizes for different displays and various v

  88. Peter Böhm, Archie C. Chapman, Pauline Pounds

    In this work we present Deep Reinforcement Learning (DRL) training of directional locomotion for low-cost quadrupedal robots in the real world. In particular, we exploit randomization of heading that the robot must follow to foster exploration of action-state transitions most useful for learning both forward locomotion as well as course adjustments. Changing

  89. Qiyin Huang, Ruomin Sui, Lunwei Zhang, Yenhang Zhou

    Grasping the same object in different postures is often necessary, especially when handling tools or stacked items. Due to unknown object properties and changes in grasping posture, the required grasping force is uncertain and variable. Traditional methods rely on real-time feedback to control the grasping force cautiously, aiming to prevent slipping or dama

  90. Kyle Sargent, Kyle Hsu, Justin Johnson, Li Fei-Fei

    Since the advent of popular visual generation frameworks like VQGAN and latent diffusion models, state-of-the-art image generation systems have generally been two-stage systems that first tokenize or compress visual data into a lower-dimensional latent space before learning a generative model. Tokenizer training typically follows a standard recipe in which i

  91. Erik Bates, Blan Morrison, Mason Rogers, Arianna Serafini

    The sequence of partial sums of Fibonacci numbers, beginning with $2$, $4$, $7$, $12$, $20$, $33,\dots$, has several combinatorial interpretations (OEIS A000071). For instance, the $n$-th term in this sequence is the number of length-$n$ binary words that avoid $110$. This paper proves a related but new interpretation: given a length-$3$ binary word -- calle

  92. Worameth Chinchuthakun, Tossaporn Saengja, Nontawat Tritrong, Pitchaporn Rewatbowornwong

    While diffusion models show promising results in image editing given a target prompt, achieving both prompt fidelity and background preservation remains difficult. Recent works have introduced score distillation techniques that leverage the rich generative prior of text-to-image diffusion models to solve this task without additional fine-tuning. However, the

  93. Yuhao Liu, Nian Yang, Gongqiu Zhang

    This paper develops general approaches for pricing various types of American-style Parisian options (down-in/-out, perpetual/finite-maturity) with general payoff functions based on continuous-time Markov chain (CTMC) approximation under general 1D time-inhomogeneous Markov models. For the down-in types, by conditioning on the Parisian stopping time, we reduc

  94. Zi-Xu Lu, Huai-Bing Zhu, Xuan Zuo, Jie Li

    Quantum magnonics based on YIG spheres provides a new arena for observing macroscopic quantum states. Here we propose to prepare two kinds of non-Gaussian magnonic states by adding a single magnon onto two Gaussian states, namely, coherent and thermal states. We adopt an optomagnonic system of a YIG sphere and use fast optical pulses to weakly activate the m

  95. Jieyi Tan, Chengwei Zhang, Bo Dang, Yansheng Li

    Traditional Remote Sensing Foundation models (RSFMs) are pre-trained with a data-centralized paradigm, through self-supervision on large-scale curated remote sensing data. For each institution, however, pre-training RSFMs with limited data in a standalone manner may lead to suboptimal performance, while aggregating remote sensing data from multiple instituti

  96. Hoang V. Tran, Khoi N. M. Nguyen, Trang Pham, Thanh T. Chu

    To overcome computational challenges of Optimal Transport (OT), several variants of Sliced Wasserstein (SW) has been developed in the literature. These approaches exploit the closed-form expression of the univariate OT by projecting measures onto (one-dimensional) lines. However, projecting measures onto low-dimensional spaces can lead to a loss of topologic

  97. Honghao Guo, Junda Huang, Ian Zhang, Boyuan Liang

    Robotic grasping and manipulation in underwater environments present unique challenges for robotic hands traditionally used on land. These challenges stem from dynamic water conditions, a wide range of object properties from soft to stiff, irregular object shapes, and varying surface frictions. One common approach involves developing finger-based hands with

  98. Lingpeng Chen, Siva Kailas, Srujan Deolasee, Wenhao Luo

    We introduce a novel distributed source seeking framework, DIAS, designed for multi-robot systems in scenarios where the number of sources is unknown and potentially exceeds the number of robots. Traditional robotic source seeking methods typically focused on directing each robot to a specific strong source and may fall short in comprehensively identifying a

  99. Jiachen Chen, Yaozu Wu, Zhen Yang, Shibo Xu

    Quantum machine learning is among the most exciting potential applications of quantum computing. However, the vulnerability of quantum information to environmental noises and the consequent high cost for realizing fault tolerance has impeded the quantum models from learning complex datasets. Here, we introduce AdaBoost.Q, a quantum adaptation of the classica

  100. Ning-Yuan Georgia Liu, Flower Yang, Mohammad S. Jalali

    Causal graphs are commonly used to understand and model complex systems. Researchers often construct these graphs from different perspectives, leading to significant variations for the same problem. Comparing causal graphs is, therefore, essential for evaluating assumptions, integrating insights, and resolving disagreements. The rise of AI tools has further