March 2025 arXiv papers — page 132
Showing 13,101–13,200 of 23,633 papers
GaussianIP: Identity-Preserving Realistic 3D Human Generation via Human-Centric Diffusion Prior
cs.CVZichen Tang, Yuan Yao, Miaomiao Cui, Liefeng Bo
Text-guided 3D human generation has advanced with the development of efficient 3D representations and 2D-lifting methods like Score Distillation Sampling (SDS). However, current methods suffer from prolonged training times and often produce results that lack fine facial and garment details. In this paper, we propose GaussianIP, an effective two-stage framewo
Indukuru Ramesh Reddy, M. Kaltak, Bongjae Kim
In this study, we present a systematic comparison of various approaches within the constrained random-phase approximation (cRPA) for calculating the Coulomb interaction parameter $U$. While defining the correlated space is straightforward for disentangled bands, the situation is more complex for entangled bands, where different projection schemes from hybrid
Bei Chen, Xiaowen Xiong, Renheng Zhang, Yitang Dai
Photonic neural networks have been considered as the promising candidates for next-generation neuromorphic computation, aiming to break both the power consumption wall and processing speed boundary of state-to-date digital computing architectures. Optics has shown its advantages in parallelism and linear manipulation. However, the lack of low-power and high-
Minding Fuzzy Regions: A Data-driven Alternating Learning Paradigm for Stable Lesion Segmentation
cs.CVLexin Fang, Yunyang Xu, Xiang Ma, Xuemei Li
Deep learning has achieved significant advancements in medical image segmentation, but existing models still face challenges in accurately segmenting lesion regions. The main reason is that some lesion regions in medical images have unclear boundaries, irregular shapes, and small tissue density differences, leading to label ambiguity. However, the existing m
A Comprehensive Characterization of Galaxy-cool CGM Connections at $z<0.4$ with DESI Year 1 Data
astro-ph.GAYu Voon Ng, Ting-Wen Lan, J. Xavier Prochaska, Amélie Saintonge
We investigate the relationships between the cool circumgalactic medium (CGM), traced by Ca II absorption lines, and galaxy properties at $z<0.4$ using $\sim900{,}000$ galaxy-quasar pairs within $200\,\rm kpc$ from the Year 1 data of the Dark Energy Spectroscopic Instrument (DESI). This large data set enables us to obtain composite spectra with sensitivity r
Yao Xu, Grace Nansamba, Anthony Skjellum, Gene Cooperman
There is new momentum behind an interoperable ABI for MPI, which will be a major component of MPI-5. This capability brings true separation of concerns to a running MPI computation. The linking and compilation of an MPI application becomes completely independent of the choice of MPI library. The MPI application is compiled once, and runs everywhere. This ABI
Mikhail Shkolnikov
Tropical caustic of a convex domain on the plane is a canonically associated tropical analytic curve inside the domain. In this note we give a graphical proof for the classification of its intermediate vertices, implying in particular that they are always trivalent. Apart from that we explain how various known examples of tropical caustics are constructed an
Yutaka Hosotani, Shuichiro Funatsu, Hisaki Hatanaka, Yuta Orikasa
In the $SO(5) \times U(1) \times SU(3)$ gauge-Higgs unification in the Randall-Sundrum warped space the mixing in the $W$ couplings in the quark sector is induced by masses of $SO(5)$ singlet fermions which are responsible for splitting masses of down-type quarks from those of up-type quarks in each generation. We show that the observed Cabibbo-Kobayashi-Mas
Hyeon Woo Park, Shu Zhang, Peter Meisenheimer, Maya Ramesh
Multiferroic materials, characterized by the occurrence of two or more ferroic properties, hold potential in future technological applications and also exhibit intriguing phenomena caused by the interplay of multiple orders. One such example is the formation of spin cycloid structures within multiferroic materials, which we investigate in this work by focusi
Manato Sakai, Yasuhiro Yamaguchi
Recently, a number of exotic hadrons have been reported in the experiments, and most of these states lie slightly below the threshold. Therefore, these states are considered to be hadronic molecules composed of mesons or baryons. In 2022, the doubly charmed tetraquark $T_{cc}$ was reported by the LHCb experiment, which is considered to be composed of two hea
SpaceSeg: A High-Precision Intelligent Perception Segmentation Method for Multi-Spacecraft On-Orbit Targets
cs.CVHao Liu, Pengyu Guo, Siyuan Yang, Zeqing Jiang
With the continuous advancement of human exploration into deep space, intelligent perception and high-precision segmentation technology for on-orbit multi-spacecraft targets have become critical factors for ensuring the success of modern space missions. However, the complex deep space environment, diverse imaging conditions, and high variability in spacecraf
Guihong Li, Mehdi Rezagholizadeh, Mingyu Yang, Vikram Appia
Multi-head latent attention (MLA) is designed to optimize KV cache memory through low-rank key-value joint compression. Rather than caching keys and values separately, MLA stores their compressed latent representations, reducing memory overhead while maintaining the performance. While MLA improves memory efficiency without compromising language model accurac
Vijay Bhattiprolu, Venkatesan Guruswami, Xuandi Ren
We give simple deterministic reductions demonstrating the NP-hardness of approximating the nearest codeword problem and minimum distance problem within arbitrary constant factors (and almost-polynomial factors assuming NP cannot be solved in quasipolynomial time). The starting point is a simple NP-hardness result without a gap, and is thus "PCP-free." Our ap
Ruojing Zhao, Yifei Xu, Songjie Yang, Hua Chen
In the development of wireless communication technology, multiple-input multiple-output (MIMO) technology has emerged as a key enabler, significantly enhancing the capacity of communication systems. However, traditional MIMO systems, which rely on fixed-position antennas (FPAs) with spacing limitations, cannot fully exploit the channel variations in the cont
Yijia Xu, Jianzhong Ju, Jian Luan, Jinshi Cui
The raster-ordered image token sequence exhibits a significant Euclidean distance between index-adjacent tokens at line breaks, making it unsuitable for autoregressive generation. To address this issue, this paper proposes Direction-Aware Diagonal Autoregressive Image Generation (DAR) method, which generates image tokens following a diagonal scanning order.
Directional Gaussian hypergeometric beta distributions and their uses in contaminated binary sampling
math.STBen O'Neill
We examine the Gaussian hypergeometric beta distribution and look at the effect of having an additional term in the density kernel relative to the standard beta distribution. We reparameterise and classify this distribution into left and right directional variants using parameters that give a simple and symmetrical representation of the directional push/pull
Matthew Khoriaty, Andrii Shportko, Gustavo Mercier, Zach Wood-Doughty
Recent developments in Large Language Model (LLM) capabilities have brought great potential but also posed new risks. For example, LLMs with knowledge of bioweapons, advanced chemistry, or cyberattacks could cause violence if placed in the wrong hands or during malfunctions. Because of their nature as near-black boxes, intuitive interpretation of LLM interna
Jie Liu, Yiwei Zhang, Yuan Sheng, Yujia Lou
This study proposes a dynamic rule data mining algorithm based on an improved Transformer architecture, aiming to improve the accuracy and efficiency of rule mining in a dynamic data environment. With the increase in data volume and complexity, traditional data mining methods are difficult to cope with dynamic data with strong temporal and variable character
Yongyi Jia, Shu Miao, Jiayu Wu, Ming Yang
While magnetic micro-robots have demonstrated significant potential across various applications, including drug delivery and microsurgery, the open issue of precise navigation and control in complex fluid environments is crucial for in vivo implementation. This paper introduces a novel flow-aware navigation and control strategy for magnetic micro-robots that
Songjie Yang, Jiahe Guo, Zilin He, Boyu Ning
Flexible-geometry arrays have garnered much attention in wireless communications, which dynamically adjust wireless channels to improve the system performance. In this paper, we propose a novel flexible-geometry array for a $360^\circ$ coverage, named flxible cylindrical array (FCLA), comprised of multiple flexible circular arrays (FCAs). The elements in eac
DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation
cs.CVHongbin Lin, Zilu Guo, Yifan Zhang, Shuaicheng Niu
In autonomous driving, vision-centric 3D detection aims to identify 3D objects from images. However, high data collection costs and diverse real-world scenarios limit the scale of training data. Once distribution shifts occur between training and test data, existing methods often suffer from performance degradation, known as Out-of-Distribution (OOD) problem
Fast spectral line calculations with the escape probability method and tests with synthetic observations of interstellar clouds
astro-ph.IMMika Juvela
Radiative transfer effects need to be taken into account when analysing spectral line observations. When the data are not sufficient for detailed modelling, simpler methods are needed. The escape probability formalism (EPF) is one such tool. We wish to quantify the model errors in the EPF analysis of interstellar clouds and cores. We introduce PEP, a paralle
A Multi-Objective Evaluation Framework for Analyzing Utility-Fairness Trade-Offs in Machine Learning Systems
cs.LGGökhan Özbulak, Oscar Jimenez-del-Toro, Maíra Fatoretto, Lilian Berton
The evaluation of fairness models in Machine Learning involves complex challenges, such as defining appropriate metrics, balancing trade-offs between utility and fairness, and there are still gaps in this stage. This work presents a novel multi-objective evaluation framework that enables the analysis of utility-fairness trade-offs in Machine Learning systems
Weifeng Shang, Jose Abel Castellanos Joo, Chenqi Mou, Deepak Kapur
New results on computing certificates of strictly positive polynomials in Archimedean quadratic modules are presented. The results build upon (i) Averkov's method for generating a strictly positive polynomial for which a membership certificate can be more easily computed than the input polynomial whose certificate is being sought, and (ii) Lasserre's method
UMB@PerAnsSumm 2025: Enhancing Perspective-Aware Summarization with Prompt Optimization and Supervised Fine-Tuning
cs.CLKristin Qi, Youxiang Zhu, Xiaohui Liang
We present our approach to the PerAnsSumm Shared Task, which involves perspective span identification and perspective-aware summarization in community question-answering (CQA) threads. For span identification, we adopt ensemble learning that integrates three transformer models through averaging to exploit individual model strengths, achieving an 82.91% F1-sc
Kaixuan Jiang, Yang Liu, Weixing Chen, Jingzhou Luo
Embodied Question Answering (EQA) is a challenging task in embodied intelligence that requires agents to dynamically explore 3D environments, actively gather visual information, and perform multi-step reasoning to answer questions. However, current EQA approaches suffer from critical limitations in exploration efficiency, dataset design, and evaluation metri
Hanbyul Song, Miguel F. Santos Silva, Jaume Suau, Luis Espinosa-Anke
Understanding why people trust or distrust one another, institutions, or information is a complex task that has led scholars from various fields of study to employ diverse epistemological and methodological approaches. Despite the challenges, it is generally agreed that the antecedents of trust (and distrust) encompass a multitude of emotional and cognitive
Jun Yu, Yunxiang Zhang, Xilong Lu, Yang Zheng
In this report, we present our solution for the Action Unit (AU) Detection Challenge, in 8th Competition on Affective Behavior Analysis in-the-wild. In order to achieve robust and accurate classification of facial action unit in the wild environment, we introduce an innovative method that leverages audio-visual multimodal data. Our method employs ConvNeXt as
Vishal Gandhi, Sagar Gandhi
The rise of large language models (LLMs) has revolutionized natural language processing (NLP), yet the influence of prompt sentiment, a latent affective characteristic of input text, remains underexplored. This study systematically examines how sentiment variations in prompts affect LLM-generated outputs in terms of coherence, factuality, and bias. Leveragin
Guillermo Nuñez Ponasso
We study the maximum absolute value of the determinant of matrices with entries in the set of $\ell$-th roots of unity; this is a generalization of $D$-optimal designs and Hadamard's maximal determinant problem, which involves $\pm 1$ matrices. For general values of $\ell$, we give sharpened determinantal upper bounds and constructions of matrices of large d
Yanwei Huang, Wesley Hanwen Deng, Sijia Xiao, Motahhare Eslami
Generative text-to-image (T2I) models are known for their risks related such as bias, offense, and misinformation. Current AI auditing methods face challenges in scalability and thoroughness, and it is even more challenging to enable auditors to explore the auditing space in a structural and effective way. Vipera employs multiple visual cues including a scen
Songjie Yang, Zihang Wan, Boyu Ning, Weidong Mei
Typical reconfigurable intelligent surface (RIS) implementations include metasurfaces with almost passive unit elements capable of reflecting their incident waves in controllable ways, enhancing wireless communications in a cost-effective manner. In this paper, we advance the concept of intelligent metasurfaces by introducing a flexible array geometry, terme
Joint Optimization of Resource Allocation and Radar Receiver Selection in Integrated Communication-Radar Systems
eess.SPChen Zhong, Xufeng Zhou, Lan Tang, Mengting Lou
In this paper, we investigate a distributed multi-input multi-output and orthogonal frequency division multiplexing (MIMO-OFDM) dual-function radar-communication (DFRC) system, which enables simultaneous communication and sensing in different subcarrier sets. To obtain the best tradeoff between communication and sensing performance, we first derive Cramer-Ra
Tuning Electrode Wettability to Optimize Nanobubble Nucleation and Reaction Rates in Electrochemical Gas-Evolving Reactions
cond-mat.softZhenlei Wanga, Yaxi Yua, Mengkai Qin, Hao Jiang
Bubble formation in electrochemical system often hinders reaction efficiency by reducing active surface area and obstructing mass transfer, yet the mechanisms governing their nanoscale nucleation dynamics and impact remains unclear. In this study, we used molecular dynamics simulations to explore nanobubble nucleation and reaction rates during water electrol
Julia Gersey, Rose Allegrette, Joshua Lian, Zawad Munshi
The growing homelessness crisis in the U.S. presents complex social, economic, and public health challenges, straining shelters, healthcare, and social services while limiting effective interventions. Traditional assessment methods struggle to capture its dynamic, dispersed nature, highlighting the need for scalable, data-driven detection. This survey explor
Microassembly of Multi-Material and 3D Integration Enabled by Programmable and Universal High-Precision Micro-Transfer Printing
physics.app-phQinhua Guo, Lizhou Yang, Yawen Gan, Jingyang Zhang
Micro-transfer printing is an assembly technology that enables large-scale integration of diverse materials and components from micro- to nano-scale. However, traditional micro-transfer printing technologies lack dynamic selectivity, limiting capabilities in sorting and repairing materials and components for effective yield management during large-scale manu
Yifang Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi
The key-value (KV) cache in the tensor version of transformers presents a significant bottleneck during inference. While previous work analyzes the fundamental space complexity barriers in standard attention mechanisms [Haris and Onak, 2025], our work generalizes the space complexity barriers result to tensor attention version. Our theoretical contributions
Song Cao, Taikun Zhu, Kai Jin
This paper addresses resource allocation problem with a separable objective function under a single linear constraint, formulated as maximizing $\sum_{j=1}^{n}R_j(x_j)$ subject to $\sum_{j=1}^{n}x_j=k$ and $x_j\in\{0,\dots,m\}$. While classical dynamic programming approach solves this problem in $O(n^2m^2)$ time, we propose a regret-enabled greedy algorithm
Jiawei Wang, Xuedong Hu, Herbert F Fotso
We study the low-energy spectrum of a single hole confined in a planar Ge quantum dot (QD) within the effective-mass formalism. The QD is sandwiched between two GeSi barriers of finite potential height grown along the [001] direction. To treat this finite barrier problem, we adopt an independent-band approach in dealing with boundary conditions. The effects
Shivam Dubey, Mrinal Kanti Roychowdhury, Saurabh Verma
For a given $r\in (0, +\infty)$, the quantization dimension of order $r$, if it exists, denoted by $D_r(μ)$, of a Borel probability measure $μ$ on ${\mathbb R}^d$ represents the speed how fast the $n$th quantization error of order $r$ approaches to zero as the number of elements $n$ in an optimal set of $n$-means for $μ$ tends to infinity. If $D_r(μ)$ does n
Lei Qin, Ye Pu
Optimization problems involving the minimization of a finite sum of smooth, possibly non-convex functions arise in numerous applications. To achieve a consensus solution over a network, distributed optimization algorithms, such as \textbf{EXTRA} (decentralized exact first-order algorithm), have been proposed to address these challenges. In this paper, we ana
Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
cs.CVAvinash Madasu, Vasudev Lal, Phillip Howard
CLIP is one of the most popular foundation models and is heavily used for many vision-language tasks, yet little is known about its inner workings. As CLIP is increasingly deployed in real-world applications, it is becoming even more critical to understand its limitations and embedded social biases to mitigate potentially harmful downstream consequences. How
Deep Learning-based OTFS Channel Estimation and Symbol Detection with Plug-and-Play Framework
eess.SPXiaoqi Zhang, Zhitong Ni, Weijie Yuan, J. Andrew Zhang
Orthogonal Time Frequency Space (OTFS) modulation has recently attracted significant interest due to its potential for enabling reliable communication in high-mobility environments. However, the effectiveness of OTFS receivers relies on the inherent characteristic of the Delay-Doppler (DD) domain channel, where the sparsity of the discretized channel varies
Asifullah Khan, Laiba Asmatullah, Anza Malik, Shahzaib Khan
Self-supervised learning is a machine learning approach that generates implicit labels by learning underlined patterns and extracting discriminative features from unlabeled data without manual labelling. Contrastive learning introduces the concept of "positive" and "negative" samples, where positive pairs (e.g., variation of the same image/object) are brough
The Role of Hydrogen and Oxygen Interstitial Defects in Crystalline Si cells: Mechanism of Device Degradation in Humid Environment
cond-mat.mtrl-sciBo Li, Feifei Zhang, Yu Pang, Jinyu Hu
The efficiency of silicon solar cells gradually decreases in various environments, with humidity being a key factor contributing to this decline through moisture-induced degradation (MID) involving multiple mechanisms including encapsulant hydrolysis and metal ion migration. Among these mechanisms, the role of water-derived hydrogen and oxygen interstitial d
Arnab Bhattacharyya, Weiming Feng, Piyush Srivastava
The total variation distance is a metric of central importance in statistics and probability theory. However, somewhat surprisingly, questions about computing it algorithmically appear not to have been systematically studied until very recently. In this paper, we contribute to this line of work by studying this question in the important special case of multi
AI-assisted hyper-dimensional broadband quantum memory with efficiency above 90% in warm atoms
quant-phZeliang Wu, Jinxian Guo, Zhifei Yu, Wenfeng Huang
High-dimensional broadband quantum memory significantly expands quantum information processing capabilities, but the memory efficiency becomes insufficient when extended to high dimensions. We demonstrate an efficient quantum memory for hyper-dimensional photons encoded with orbital angular momentum (OAM) and spin angular momentum (SAM). OAM information is e
Wenbang Deng, Xieyuanli Chen, Qinghua Yu, Yunze He
Semantic segmentation is a key technique that enables mobile robots to understand and navigate surrounding environments autonomously. However, most existing works focus on segmenting known objects, overlooking the identification of unknown classes, which is common in real-world applications. In this paper, we propose a feature-oriented framework for open-set
Augmenting Image Annotation: A Human-LMM Collaborative Framework for Efficient Object Selection and Label Generation
cs.CVHe Zhang, Xinyi Fu, John M. Carroll
Traditional image annotation tasks rely heavily on human effort for object selection and label assignment, making the process time-consuming and prone to decreased efficiency as annotators experience fatigue after extensive work. This paper introduces a novel framework that leverages the visual understanding capabilities of large multimodal models (LMMs), pa
Stationary solutions to the critical and super-critical quasi-geostrophic equation in the scaling critical Sobolev space
math.APMikihiro Fujii
We consider the stationary problem for the quasi-geostrophic equation with the critical and super-critical dissipation and prove the unique existence of small solutions for given small external force in the scaling critical Sobolev spaces framework. Moreover, we also show that the data-to-solution map is continuous. Since the critical and super-critical case
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
cs.CVWeichen Zhang, Zile Zhou, Xin Zeng, Xuchen Liu
Spatial reasoning is a fundamental capability of multimodal large language models (MLLMs), yet their performance in open aerial environments remains underexplored. In this work, we present Open3D-VQA, a novel benchmark for evaluating MLLMs' ability to reason about complex spatial relationships from an aerial perspective. The benchmark comprises 73k QA pairs
Yuan Liu, Saihui Hou, Saijie Hou, Jiabao Du
Image Difference Captioning (IDC) aims to generate natural language descriptions of subtle differences between image pairs, requiring both precise visual change localization and coherent semantic expression. Despite recent advancements, existing datasets often lack breadth and depth, limiting their applicability in complex and dynamic environments: (1) from
Mikihiro Fujii, Tsukasa Iwabuchi
We consider the stationary problem for the quasi-geostrophic equation on the whole plane and investigate its well-posedness and ill-posedness. In[Fujii, Ann. PDE 10, 10 (2024)], it was shown that the two-dimensional stationary Navier--Stokes equations are ill-posed in the critical Besov spaces $\dot B_{p,1}^{\frac{2}{p}-1}(\mathbb{R}^2)$ with $1 \leq p \leq
Ganlong Zhao, Guanbin Li, Jia Pan, Yizhou Yu
Aerial Vision-and-Language Navigation (Aerial VLN) aims to obtain an unmanned aerial vehicle agent to navigate aerial 3D environments following human instruction. Compared to ground-based VLN, aerial VLN requires the agent to decide the next action in both horizontal and vertical directions based on the first-person view observations. Previous methods strugg
Sustainable Grid through Distributed Data Centers: Spinning AI Demand for Grid Stabilization and Optimization
cs.DCScott C Evans, Nathan Dahlin, Ibrahima Ndiaye, Sachini Piyoni Ekanayake
We propose a disruptive paradigm to actively place and schedule TWhrs of parallel AI jobs strategically on the grid, at distributed, grid-aware high performance compute data centers (HPC) capable of using their massive power and energy load to stabilize the grid while reducing grid build-out requirements, maximizing use of renewable energy, and reducing Gree
Machine-learning heat flux closure for multi-moment fluid modeling of nonlinear Landau damping
physics.plasm-phZiyu Huang, Chuanfei Dong, Liang Wang
Nonlinear plasma physics problems are usually simulated through comprehensive modeling of phase space. The extreme computational cost of such simulations has motivated the development of multi-moment fluid models. However, a major challenge has been finding a suitable fluid closure for these fluid models. Recent developments in physics-informed machine learn
Yi Zhang, Qiang Zhang, Xiaozhu Ju, Zhaoyang Liu
While multimodal large language models (MLLMs) have made groundbreaking progress in embodied intelligence, they still face significant challenges in spatial reasoning for complex long-horizon tasks. To address this gap, we propose EmbodiedVSR (Embodied Visual Spatial Reasoning), a novel framework that integrates dynamic scene graph-guided Chain-of-Thought (C
Yifan Liu, Xun Xu, Shijie Li, Jingyi Liao
Multi-camera systems provide richer contextual information for industrial anomaly detection. However, traditional methods process each view independently, disregarding the complementary information across viewpoints. Existing multi-view anomaly detection approaches typically employ data-driven cross-view attention for feature fusion but fail to leverage the
Stark difference in the in-plane anomalous Hall response in Zintl compounds EuA2Sb2 (A = Zn, Cd) thin films
cond-mat.mtrl-sciHsiang Lee, Shinichi Nishihaya, Markus Kriener, Jun Fujioka
Recent observation of the in-plane anomalous Hall effect in magnetic Weyl semimetal EuCd2Sb2 has drawn attention to out-of-plane orbital magnetization induced by an in-plane field component. Here we study EuZn2Sb2, a sister compound of EuCd2Sb2, to demonstrate sensitive changes of the in-plane anomalous Hall effect on the band modulation. The Hall resistivit
Haihong Zhao, Zhixun Li, Chenyi Zi, Aochuan Chen
Graph learning plays a vital role in mining and analyzing complex relationships within graph data and has been widely applied to real-world scenarios such as social, citation, and e-commerce networks. Foundation models in computer vision (CV) and natural language processing (NLP) have demonstrated remarkable cross-domain capabilities that are equally signifi
Sixiang Ye, Zeyu Sun, Guoqing Wang, Liwei Guo
Code generation has emerged as a key task to automate software development by converting high-level descriptions into executable code. Large language models (LLMs) excel at this but depend heavily on input prompt quality.Manual prompt engineering can be time-consuming and inconsistent, limiting LLM effectiveness. This paper introduces Prochemy, an innovative
Zhou Fang, Hanlu Zhang, Jacky He, Zhen Qi
This study aims to develop an efficient and accurate model for detecting malicious comments, addressing the increasingly severe issue of false and harmful content on social media platforms. We propose a deep learning model that combines BERT and BiLSTM. The BERT model, through pre-training, captures deep semantic features of text, while the BiLSTM network ex
Yangyang Xie, Cheng Hu, Nicolas Baumann, Edoardo Ghignone
Autonomous drifting is a complex challenge due to the highly nonlinear dynamics and the need for precise real-time control, especially in uncertain environments. To address these limitations, this paper presents a hierarchical control framework for autonomous vehicles drifting along general paths, primarily focusing on addressing model inaccuracies and mitig
Liwei Guo, Sixiang Ye, Zeyu Sun, Xiang Chen
Large Language Models (LLMs) have demonstrated remarkable performance in code completion. However, the training data used to develop these models often contain a significant amount of buggy code. Yet, it remains unclear to what extent these buggy instances influence LLMs' performance when tackling bug-prone code completion tasks. To fill this gap, this paper
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
cs.ROPingrui Zhang, Xianqiang Gao, Yuhan Wu, Kehui Liu
In mobile manipulation, navigation and manipulation are often treated as separate problems, resulting in a significant gap between merely approaching an object and engaging with it effectively. Many navigation approaches primarily define success by proximity to the target, often overlooking the necessity for optimal positioning that facilitates subsequent ma
Wuwei Huang, Renren Jin, Wen Zhang, Jian Luan
Recent studies on end-to-end speech translation(ST) have facilitated the exploration of multilingual end-to-end ST and end-to-end simultaneous ST. In this paper, we investigate end-to-end simultaneous speech translation in a one-to-many multilingual setting which is closer to applications in real scenarios. We explore a separate decoder architecture and a un
Zhiyu Dong, Patrick A. Lee
Can strong repulsive interactions be shown to give rise to pairing in a controlled way? We find that for a single flavor polarized band, there is a small expansion parameter in the low density limit, once the Bloch wavefunction form factor is taken into account. A perturbative expansion is possible, even if the interaction is much stronger than the Fermi ene
Taehwan Lee, Kyeongkook Seo, Jaejun Yoo, Sung Whan Yoon
Flat minima, known to enhance generalization and robustness in supervised learning, remain largely unexplored in generative models. In this work, we systematically investigate the role of loss surface flatness in generative models, both theoretically and empirically, with a particular focus on diffusion models. We establish a theoretical claim that flatter m
Joseph Oglio, Mikhail Nesterenko, Gokarna Sharma
We present SmartShards: a new sharding algorithm for improving Byzantine tolerance and churn resistance in blockchains. Our algorithm places a peer in multiple shards to create an overlap. This simplifies cross-shard communication and shard membership management. We describe SmartShards, prove it correct and evaluate its performance. We propose several Smart
Pinghui Huang, Xue-Ning Bai
Dust concentration in protoplanetary disks (PPDs) is the first step towards planetesimal formation, a crucial yet highly uncertain stage in planet formation. Although the streaming instability (SI) is widely recognized as a powerful mechanism for planetesimal formation, its properties can be sensitive to the gas dynamical environment. The outer region of PPD
Zixiao Ma, Baosen Zhang
The growing integration of inverter-based resources (IBRs) into modern power systems poses significant challenges for maintaining reliable operation under dynamic and constrained conditions. This paper focuses on the power tracking problem for grid-connected IBRs, addressing the complexities introduced by voltage and power factor constraints. Voltage constra
Xueyang Zhou, Guiyao Tie, Guowen Zhang, Weidong Wang
The rise of Large Reasoning Models (LRMs) signifies a paradigm shift toward advanced computational reasoning. Yet, this progress disrupts traditional agent frameworks, traditionally anchored by execution-oriented Large Language Models (LLMs). To explore this transformation, we propose the LaRMA framework, encompassing nine tasks across Tool Usage, Plan Desig
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
cs.CVHongyang Wei, Shuaizheng Liu, Chun Yuan, Lei Zhang
By leveraging the generative priors from pre-trained text-to-image diffusion models, significant progress has been made in real-world image super-resolution (Real-ISR). However, these methods tend to generate inaccurate and unnatural reconstructions in complex and/or heavily degraded scenes, primarily due to their limited perception and understanding capabil
Haotian Tan, Yuan-Hua Ni
Time-optimal trajectory planning and control is central for autonomous vehicles, yet its application and real-time deployment confronts two fundamental challenges: the non-convexity of optimal control problems and the unpredictable computation time inherent to nonlinear programming. To address these challenges, we propose a hierarchical convex optimization f
Zhenguang Liu, Chao Shuai, Shaojing Fan, Ziping Dong
Diffusion models have achieved remarkable success in novel view synthesis, but their reliance on large, diverse, and often untraceable Web datasets has raised pressing concerns about image copyright protection. Current methods fall short in reliably identifying unauthorized image use, as they struggle to generalize across varied generation tasks and fail whe
Kelu Yao, Nuo Xu, Rong Yang, Yingying Xu
This paper introduces a holistic vision-language foundation model tailored for remote sensing, named Falcon. Falcon offers a unified, prompt-based paradigm that effectively executes comprehensive and complex remote sensing tasks. Falcon demonstrates powerful understanding and reasoning abilities at the image, region, and pixel levels. Specifically, given sim
Hyunwoo Park, Baekryun Seong, Sang-Ki Ko
In cooperative multi-agent reinforcement learning (MARL), the permutation problem where the state space grows exponentially with the number of agents reduces sample efficiency. Additionally, many existing architectures struggle with scalability, relying on a fixed structure tied to a specific number of agents, limiting their applicability to environments wit
Chaoyun Zhang, Shilin He, Liqun Li, Si Qin
Large language models (LLMs) have evolved beyond simple text generation to power software agents that directly translate natural language commands into tangible actions. While API-based LLM agents initially rose to prominence for their robust automation capabilities and seamless integration with programmatic endpoints, recent progress in multimodal LLM resea
Leqi Lin, Xingyu Zhou, Kaiyuan Yang, Xizhong Chen
Pharmaceutical process design and development for generic, innovative, or personalized drugs have always been a time-consuming, costly, rigorous process, that involves multi-stage evaluation for better quality control and assurance. Large language models (LLMs), a type of generative artificial intelligence system, can augment laboratory research in the pharm
Bin Liu, Xiaohong Liu, Qin Luo, Ziqiao Shang
Pairwise learning underpins implicit collaborative filtering, yet its effectiveness is often hindered by sparse supervision, noisy interactions, and popularity-driven exposure bias. In this paper, we propose Variational Bayesian Personalized Ranking (VarBPR), a tractable variational framework for implicit-feedback pairwise learning that offers principled exp
Further exploration of binding energy residuals using machine learning and the development of a composite ensemble model
nucl-thI. Bentley, J. Tedder, M. Gebran, A. Paul
This paper describes the development of the Four Model Tree Ensemble (FMTE). The FMTE is a composite of machine learning models trained on experimental binding energies from the Atomic Mass Evaluation (AME) 2012. The FMTE predicts binding energy values for all nuclei with N > 7 and Z > 7 from AME 2020 with a standard deviation of 76 keV and a mean average de
Low-cost Real-world Implementation of the Swing-up Pendulum for Deep Reinforcement Learning Experiments
cs.LGPeter Böhm, Pauline Pounds, Archie C. Chapman
Deep reinforcement learning (DRL) has had success in virtual and simulated domains, but due to key differences between simulated and real-world environments, DRL-trained policies have had limited success in real-world applications. To assist researchers to bridge the \textit{sim-to-real gap}, in this paper, we describe a low-cost physical inverted pendulum a
MobiVital: Self-supervised Time-series Quality Estimation for Contactless Respiration Monitoring Using UWB Radar
eess.SPZiqi Wang, Derek Hua, Wenjun Jiang, Tianwei Xing
Respiration waveforms are increasingly recognized as important biomarkers, offering insights beyond simple respiration rates, such as detecting breathing irregularities for disease diagnosis or monitoring breath patterns to guide rehabilitation training. Previous works in wireless respiration monitoring have primarily focused on estimating respiration rate,
Magnetoconductivity due to electron-electron interaction in a non-Galilean-invariant Fermi liquid
cond-mat.str-elTatia Kiliptari, Dmitrii L. Maslov
The $T^2$-scaling of resistivity with temperature is often viewed as a classic hallmark of a Fermi-liquid (FL) behavior in metals. However, if umklapp scattering is suppressed, this scaling is not universally guaranteed to occur. In this case, the resistivity behavior is influenced by several factors, such as dimensionality (two vs. three), topology (simply-
Wenhao Jiang, Duo Li, Menghan Hu, Chao Ma
In the field of autonomous driving, end-to-end deep learning models show great potential by learning driving decisions directly from sensor data. However, training these models requires large amounts of labeled data, which is time-consuming and expensive. Considering that the real-world driving data exhibits a long-tailed distribution where simple scenarios
Jordan S. Ellenberg, Cristofero S. Fraser-Taliente, Thomas R. Harvey, Karan Srivastava
We present a new implementation of the LLM-driven genetic algorithm {\it funsearch}, whose aim is to generate examples of interest to mathematicians and which has already had some success in problems in extremal combinatorics. Our implementation is designed to be useful in practice for working mathematicians; it does not require expertise in machine learning
Heng Wang, Yotaro Shimose, Shingo Takamatsu
Advertising banners are critical for capturing user attention and enhancing advertising campaign effectiveness. Creating aesthetically pleasing banner designs while conveying the campaign messages is challenging due to the large search space involving multiple design elements. Additionally, advertisers need multiple sizes for different displays and various v
Training Directional Locomotion for Quadrupedal Low-Cost Robotic Systems via Deep Reinforcement Learning
cs.ROPeter Böhm, Archie C. Chapman, Pauline Pounds
In this work we present Deep Reinforcement Learning (DRL) training of directional locomotion for low-cost quadrupedal robots in the real world. In particular, we exploit randomization of heading that the robot must follow to foster exploration of action-state transitions most useful for learning both forward locomotion as well as course adjustments. Changing
Qiyin Huang, Ruomin Sui, Lunwei Zhang, Yenhang Zhou
Grasping the same object in different postures is often necessary, especially when handling tools or stacked items. Due to unknown object properties and changes in grasping posture, the required grasping force is uncertain and variable. Traditional methods rely on real-time feedback to control the grasping force cautiously, aiming to prevent slipping or dama
Kyle Sargent, Kyle Hsu, Justin Johnson, Li Fei-Fei
Since the advent of popular visual generation frameworks like VQGAN and latent diffusion models, state-of-the-art image generation systems have generally been two-stage systems that first tokenize or compress visual data into a lower-dimensional latent space before learning a generative model. Tokenizer training typically follows a standard recipe in which i
Erik Bates, Blan Morrison, Mason Rogers, Arianna Serafini
The sequence of partial sums of Fibonacci numbers, beginning with $2$, $4$, $7$, $12$, $20$, $33,\dots$, has several combinatorial interpretations (OEIS A000071). For instance, the $n$-th term in this sequence is the number of length-$n$ binary words that avoid $110$. This paper proves a related but new interpretation: given a length-$3$ binary word -- calle
Worameth Chinchuthakun, Tossaporn Saengja, Nontawat Tritrong, Pitchaporn Rewatbowornwong
While diffusion models show promising results in image editing given a target prompt, achieving both prompt fidelity and background preservation remains difficult. Recent works have introduced score distillation techniques that leverage the rich generative prior of text-to-image diffusion models to solve this task without additional fine-tuning. However, the
Yuhao Liu, Nian Yang, Gongqiu Zhang
This paper develops general approaches for pricing various types of American-style Parisian options (down-in/-out, perpetual/finite-maturity) with general payoff functions based on continuous-time Markov chain (CTMC) approximation under general 1D time-inhomogeneous Markov models. For the down-in types, by conditioning on the Parisian stopping time, we reduc
Zi-Xu Lu, Huai-Bing Zhu, Xuan Zuo, Jie Li
Quantum magnonics based on YIG spheres provides a new arena for observing macroscopic quantum states. Here we propose to prepare two kinds of non-Gaussian magnonic states by adding a single magnon onto two Gaussian states, namely, coherent and thermal states. We adopt an optomagnonic system of a YIG sphere and use fast optical pulses to weakly activate the m
Towards Privacy-preserved Pre-training of Remote Sensing Foundation Models with Federated Mutual-guidance Learning
cs.CVJieyi Tan, Chengwei Zhang, Bo Dang, Yansheng Li
Traditional Remote Sensing Foundation models (RSFMs) are pre-trained with a data-centralized paradigm, through self-supervision on large-scale curated remote sensing data. For each institution, however, pre-training RSFMs with limited data in a standalone manner may lead to suboptimal performance, while aggregating remote sensing data from multiple instituti
Hoang V. Tran, Khoi N. M. Nguyen, Trang Pham, Thanh T. Chu
To overcome computational challenges of Optimal Transport (OT), several variants of Sliced Wasserstein (SW) has been developed in the literature. These approaches exploit the closed-form expression of the univariate OT by projecting measures onto (one-dimensional) lines. However, projecting measures onto low-dimensional spaces can lead to a loss of topologic
Honghao Guo, Junda Huang, Ian Zhang, Boyuan Liang
Robotic grasping and manipulation in underwater environments present unique challenges for robotic hands traditionally used on land. These challenges stem from dynamic water conditions, a wide range of object properties from soft to stiff, irregular object shapes, and varying surface frictions. One common approach involves developing finger-based hands with
Lingpeng Chen, Siva Kailas, Srujan Deolasee, Wenhao Luo
We introduce a novel distributed source seeking framework, DIAS, designed for multi-robot systems in scenarios where the number of sources is unknown and potentially exceeds the number of robots. Traditional robotic source seeking methods typically focused on directing each robot to a specific strong source and may fall short in comprehensively identifying a
Jiachen Chen, Yaozu Wu, Zhen Yang, Shibo Xu
Quantum machine learning is among the most exciting potential applications of quantum computing. However, the vulnerability of quantum information to environmental noises and the consequent high cost for realizing fault tolerance has impeded the quantum models from learning complex datasets. Here, we introduce AdaBoost.Q, a quantum adaptation of the classica
Ning-Yuan Georgia Liu, Flower Yang, Mohammad S. Jalali
Causal graphs are commonly used to understand and model complex systems. Researchers often construct these graphs from different perspectives, leading to significant variations for the same problem. Comparing causal graphs is, therefore, essential for evaluating assumptions, integrating insights, and resolving disagreements. The rise of AI tools has further