Skip to content

May 2025 arXiv papers — page 125

Showing 12,40112,500 of 24,552 papers

  1. Yinzhe Wang, Yiwen Xiao, Hu Wang, Yiping Xu

    Multi-view stereo (MVS) models based on progressive depth hypothesis narrowing have made remarkable advancements. However, existing methods haven't fully utilized the potential that the depth coverage of individual instances is smaller than that of the entire scene, which restricts further improvements in depth estimation precision. Moreover, inevitable devi

  2. Subhayan Saha, Giovanni Barbarino, Nicolas Gillis

    Tensor decompositions have become a central tool in data science, with applications in areas such as data analysis, signal processing, and machine learning. A key property of many tensor decompositions, such as the canonical polyadic decomposition, is identifiability: the factors are unique, up to trivial scaling and permutation ambiguities. This allows one

  3. Shogo Kusano, Masayuki Uchida

    We study structural equation modeling (SEM) for diffusion processes with jumps. Based on high-frequency data, we consider the parameter estimation and the goodness-of-fit test in the SEM. Using a threshold method, we propose the quasi-likelihood of the SEM and prove that the quasi-maximum likelihood estimator has consistency and asymptotic normality. To exam

  4. Qichen Sun, Zhengrui Guo, Rui Peng, Hao Chen

    Recent advances in computational pathology and artificial intelligence have significantly enhanced the utilization of gigapixel whole-slide images and and additional modalities (e.g., genomics) for pathological diagnosis. Although deep learning has demonstrated strong potential in pathology, several key challenges persist: (1) fusing heterogeneous data types

  5. Yingkai Kang, Jiawen Kang, Jinbo Wen, Tao Zhang

    Vehicular metaverses are an emerging paradigm that merges intelligent transportation systems with virtual spaces, leveraging advanced digital twin and Artificial Intelligence (AI) technologies to seamlessly integrate vehicles, users, and digital environments. In this paradigm, vehicular AI agents are endowed with environment perception, decision-making, and

  6. Serge Dolgikh

    This study explores the emergence of counter-inferential behavior in natural and artificial cognitive systems, that is, patterns in which agents misattribute empirical success or suppress adaptation, leading to epistemic rigidity or maladaptive stability. We analyze archetypal scenarios in which such behavior arises: reinforcement of stability through reward

  7. Zhichen Zeng, Ruizhong Qiu, Wenxuan Bao, Tianxin Wei

    Graph neural networks, despite their impressive performance, are highly vulnerable to distribution shifts on graphs. Existing graph domain adaptation (graph DA) methods often implicitly assume a mild shift between source and target graphs, limiting their applicability to real-world scenarios with large shifts. Gradual domain adaptation (GDA) has emerged as a

  8. Yingchen He, Christian D. Weilbach, Martyna E. Wojciechowska, Yuxuan Zhang

    Advances in deep generative modeling have made it increasingly plausible to train human-level embodied agents. Yet progress has been limited by the absence of large-scale, real-time, multi-modal, and socially interactive datasets that reflect the sensory-motor complexity of natural environments. To address this, we present PLAICraft, a novel data collection

  9. Elvira Di Nardo, Giuseppe Guarino

    This paper develops new combinatorial approaches to analyze and compute special set partitions, called complementary set partitions, which are fundamental in the study of generalized cumulants. Moving away from traditional graph-based and algebraic methods, a simple and fast algorithm is proposed to list complementary set partitions based on two-block partit

  10. Joel Jang, Seonghyeon Ye, Zongyu Lin, Jiannan Xiang

    We introduce DreamGen, a simple yet highly effective 4-stage pipeline for training robot policies that generalize across behaviors and environments through neural trajectories - synthetic robot data generated from video world models. DreamGen leverages state-of-the-art image-to-video generative models, adapting them to the target robot embodiment to produce

  11. Juntaro Fujii, Kazuki Yamamoto, Akihisa Koga

    We investigate an SU(3) Fermi-Hubbard model on a hypercubic lattice at finite temperatures, combining dynamical mean-field theory with continuous-time quantum Monte Carlo simulations. Taking strong correlations into account carefully, we find a ferromagnetically ordered state, in which one of the three components becomes dominant, when holes are doped away f

  12. Jiabin Chen, Haiping Wang, Jinpeng Li, Yuan Liu

    We propose SpatialLLM, a novel approach advancing spatial intelligence tasks in complex urban scenes. Unlike previous methods requiring geographic analysis tools or domain expertise, SpatialLLM is a unified language model directly addressing various spatial intelligence tasks without any training, fine-tuning, or expert intervention. The core of SpatialLLM l

  13. Tianming Liang, Haichao Jiang, Yuting Yang, Chaolei Tan

    Referring video object segmentation (RVOS) aims to identify, track and segment the objects in a video based on language descriptions, which has received great attention in recent years. However, existing datasets remain focus on short video clips within several seconds, with salient objects visible in most frames. To advance the task towards more practical s

  14. Ke Yang, Kevin Ros, Shankar Kumar Senthil Kumar, ChengXiang Zhai

    Just-in-time Information Recommendation (JIR) is a service designed to deliver the most relevant information precisely when users need it, , addressing their knowledge gaps with minimal effort and boosting decision-making and efficiency in daily life. Advances in device-efficient deployment of foundation models and the growing use of intelligent wearable dev

  15. Wanshan Cui, Yejin Jeong, Inwook Song, Gyuri Kim

    Accurate prediction of polymer material properties through data-driven approaches greatly accelerates novel material development by reducing redundant experiments and trial-and-error processes. However, inevitable outliers in empirical measurements can severely skew machine learning results, leading to erroneous prediction models and suboptimal material desi

  16. Shuyang Dong, Shangtong Zhang, Lu Feng

    Reinforcement Learning (RL) has shown great promise in domains like healthcare and robotics but often struggles with adoption due to its lack of interpretability. Counterfactual explanations, which address "what if" scenarios, provide a promising avenue for understanding RL decisions but remain underexplored for continuous action spaces. We propose a novel a

  17. Utsav Banerjee, Chiraag Juvekar, Yong Ki Lee, Leibo Liu

    Security is increasingly more important in designing chips and systems based on them, and the International Solid-State Circuits Conference (ISSCC), the leading conference for presenting advances in solid-state circuits and semiconductor technology, is committed to hardware security by establishing the security subcommittee since 2024. In the past two years,

  18. Sushmita Gupta, Pallavi Jain, Souvik Saha, Saket Saurabh

    Multiwinner Elections have emerged as a prominent area of research with numerous practical applications. We contribute to this area by designing parameterized approximation algorithms and also resolving an open question by Yang and Wang [AAMAS'18]. More formally, given a set of candidates, \mathcal{C}, a set of voters,\mathcal{V}, approving a subset of candi

  19. Yixin Chen, Xiaoyang Wang, Wanghui Li, Mohan Chen

    Tin (Sn) plays a crucial role in studying the dynamic mechanical responses of ductile metals under shock loading. Atomistic simulations serves to unveil the nano-scale mechanisms for critical behaviors of dynamic responses. However, existing empirical potentials for Sn often lack sufficient accuracy when applied in such simulation. Particularly, the solid-so

  20. Chaofan Li, Jianlyu Chen, Yingxia Shao, Defu Lian

    Code embedding models attract increasing attention due to the widespread popularity of retrieval-augmented generation (RAG) in software development. These models are expected to capture the rich semantic relationships inherent to code, which differ significantly from those found in text. However, existing models remain severely limited due to the scarcity of

  21. Wenqi Tong, H. Alaeian, F. Robicheaux

    We study the steady-state behavior of the open Dicke model, which describes the collective interaction of $N$ spin-$1/2$ particles with a lossy, quantized cavity mode and exhibits a superradiant phase transition above a critical light-matter coupling. While the standard model conserves total spin, Kirton and Keeling \cite{PhysRevLett.118.123602} demonstrated

  22. Wei Hu, Danyang Huang, Bo Zhang

    Social network platforms today generate vast amounts of data, including network structures and a large number of user-defined tags, which reflect users' interests. The dimensionality of these personalized tags can be ultra-high, posing challenges for model analysis in targeted preference analysis. Traditional categorical feature screening methods overlook th

  23. Kenya Abe, Kunihiro Takeoka, Makoto P. Kato, Masafumi Oyamada

    Query expansion (QE) enhances retrieval by incorporating relevant terms, with large language models (LLMs) offering an effective alternative to traditional rule-based and statistical methods. However, LLM-based QE suffers from a fundamental limitation: it often fails to generate relevant knowledge, degrading search performance. Prior studies have focused on

  24. Luyao Lei, Shuo Xu, Yifan Bai, Xing Wei

    The performance of multi-modal 3D occupancy prediction is limited by ineffective fusion, mainly due to geometry-semantics mismatch from fixed fusion strategies and surface detail loss caused by sparse, noisy annotations. The mismatch stems from the heterogeneous scale and distribution of point cloud and image features, leading to biased matching under fixed

  25. Khang Nguyen, Khai Nguyen, An T. Le, Jan Peters

    Robot learning in high-dimensional control settings, such as humanoid locomotion, presents persistent challenges for reinforcement learning (RL) algorithms due to unstable dynamics, complex contact interactions, and sensitivity to distributional shifts during training. Model-based methods, \textit{e.g.}, Temporal-Difference Model Predictive Control (TD-MPC),

  26. Ziwei Xu, Udit Sanghi, Mohan Kankanhalli

    Large Language Models (LLMs) are increasingly deployed in interactions where they are prompted to adopt personas. This paper investigates whether such persona conditioning affects model safety under bullying, an adversarial manipulation that applies psychological pressures in order to force the victim to comply to the attacker. We introduce a simulation fram

  27. Haojie Hou, Yan-Xia Ren, Renming Song

    Let $\{(X_t)_{t\geq 0}, \mathbb{P}_{\delta_x}, x\in E\}$ be a supercritical branching Markov process (which is not necessary symmetric) on a locally compact metric measure space $(E,\mu)$ with spatially dependent local branching mechanism. Under some assumptions on the semigroup of the spatial motion, we first prove law of iterated logarithm type results for

  28. Kian Kai Ang, Guy Farrelly, Cheryl Pope, Damith C. Ranasinghe

    We develop QUICtester, an automated approach for uncovering non-compliant behaviors in the ratified QUIC protocol implementations (RFC 9000/9001). QUICtester leverages active automata learning to abstract the behavior of a QUIC implementation into a finite state machine (FSM) representation. Unlike prior noncompliance checking methods, to help uncover state

  29. Xing-Peng Yang, Kun Xu, Zhi-Fu Gao, Long Jiang

    In the bulge of M31, the Chandra observations discovered a possible black hole (BH) ultracompact X-ray binary (UCXB) Seq.1 with an orbital period of 7.7 minutes and a maximum X-ray luminosity $L_{\rm X}=1.09^{+0.02}_{-0.01}\times10^{38}~ \rm erg\,s^{-1}$ in the $0.5-8$ keV band. The minimum orbital period of the BH UCXBs predicted by the standard magnetic br

  30. Arjun Ramesh Kaushik, Bharat Chandra Yalavarthi, Arun Ross, Vishnu Boddeti

    In today's data-driven analytics landscape, deep learning has become a powerful tool, with latent representations, known as embeddings, playing a central role in several applications. In the face analytics domain, such embeddings are commonly used for biometric recognition (e.g., face identification). However, these embeddings, or templates, can inadvertentl

  31. Li Lai, Jia Li

    The Chowla--Milnor conjecture predicts the linear independence of certain Hurwitz zeta values. In this paper, we prove that for any fixed integer $k \geqslant 2$, the dimension of the $\mathbb{Q}$-linear span of $\zeta(k,a/q)-(-1)^{k}\zeta(k,1-a/q)$ ($1 \leqslant a < q/2$, $\gcd(a,q)=1$) is at least $(c -o(1)) \cdot \log q$ as the positive integer $q \to +\i

  32. Seungmin Kim, Sohee Park, Donghyun Kim, Jisu Lee

    With the advancement of AI-based speech synthesis technologies such as Deep Voice, there is an increasing risk of voice spoofing attacks, including voice phishing and fake news, through unauthorized use of others' voices. Existing defenses that inject adversarial perturbations directly into audio signals have limited effectiveness, as these perturbations can

  33. Fei Xie, Jiahao Nie, Yujin Tang, Wenkang Zhang

    Recent State Space Models (SSM), especially Mamba, have demonstrated impressive performance in visual modeling and possess superior model efficiency. However, the application of Mamba to visual tasks suffers inferior performance due to three main constraints existing in the sequential model: 1) Casual computing is incapable of accessing global context; 2) Lo

  34. Dmitry Nesterov

    We present an alternative formulation of generalized unimodular gravity (GUMG), a class of modifications to general relativity characterized by a special partial breaking of general coordinate covariance. The action for this formulation is derived constructively through a sequence of equivalent representations, starting from the original GUMG setup and exten

  35. Yinlin Zhu, Xunkai Li, Jishuo Jia, Miao Hu

    Recent advances in graph machine learning have shifted to data-centric paradigms, driven by two emerging fields: (1) Federated graph learning (FGL) enables multi-client collaboration but faces challenges from data and task heterogeneity, limiting its practicality; (2) Graph foundation models (GFM) offer strong domain generalization but are usually trained on

  36. Yihong Huang, Chen Chu

    Key feature fields need bigger embedding dimensionality, others need smaller. This demands automated dimension allocation. Existing approaches, such as pruning or Neural Architecture Search (NAS), require training a memory-intensive SuperNet that enumerates all possible dimension combinations, which is infeasible for large feature spaces. We propose DimGrow,

  37. Hana Satou, Alan Mitkiy

    Transfer learning across domains with distribution shift remains a fundamental challenge in building robust and adaptable machine learning systems. While adversarial perturbations are traditionally viewed as threats that expose model vulnerabilities, recent studies suggest that they can also serve as constructive tools for data augmentation. In this work, we

  38. Haoyu Zhao, Yihan Geng, Shange Tang, Yong Lin

    LLM-based formal proof assistants (e.g., in Lean) hold great promise for automating mathematical discovery. But beyond syntactic correctness, do these systems truly understand mathematical structure as humans do? We investigate this question in context of mathematical inequalities -- specifically the prover's ability to recognize that the given problem simpl

  39. Zhuoheng Wang, Jinyin Zhou, Qi Wu

    Humanoid soccer dribbling is a highly challenging task that demands dexterous ball manipulation while maintaining dynamic balance. Traditional rule-based methods often struggle to achieve accurate ball control due to their reliance on fixed walking patterns and limited adaptability to real-time ball dynamics. To address these challenges, we propose a two-sta

  40. Anson Chen

    The tensions between cosmological parameter measurements from the early-universe and the late-universe datasets offer an exciting opportunity to explore new physics, if not accounted for unknown systematics. Apart from the well-known Hubble tension, a tension up to $4.9 \sigma$ in the cosmic dipole has also been reported. While the cosmic dipole is mainly in

  41. Shristi Das Biswas, Arani Roy, Kaushik Roy

    As Text-to-Image models continue to evolve, so does the risk of generating unsafe, copyrighted, or privacy-violating content. Existing safety interventions - ranging from training data curation and model fine-tuning to inference-time filtering and guidance - often suffer from incomplete concept removal, susceptibility to jail-breaking, computational ineffici

  42. Xingyu Wang, Mingsen Wang, Wenbo Shen, Rui Chang

    As the default package manager for Node.js, npm has become one of the largest package management systems in the world. To facilitate dependency management for developers, npm supports a special type of dependency, Peer Dependency, whose installation and usage differ from regular dependencies. However, conflicts between peer dependencies can trap the npm clie

  43. Won-Young Hwang, Kicheon Kang

    We propose an experimental scheme to probe the quantum statistics of two identical particles. The transition between the quantum and classical statistics of two identical particles is described by the particles having identical multiple internal energy levels. We show that effective distinguishability emerges as the thermal energy increases with respect to t

  44. Mingyuan Zhou, Yi Gu, Zhendong Wang

    Diffusion distillation has emerged as a promising strategy for accelerating text-to-image (T2I) diffusion models by distilling a pretrained score network into a one- or few-step generator. While existing methods have made notable progress, they often rely on real or teacher-synthesized images to perform well when distilling high-resolution T2I diffusion mode

  45. Zhi-Ming Yang, Huan Li

    Nonsymmorphic symmetries can give rise to Dirac semimetal (DSM) states. However, few studies have been conducted on DSMs in interacting systems. Here, we induce interacting DSM states in nonsymmorphic iridium oxides SrIrO$_3$, BaIrO$_3$ and CaIrO$_3$, and contend that the interaction of electron-electron correlations, strong spin-orbital coupling, and symmet

  46. Pengxin Guo, Yinong Wang, Wei Li, Mengting Liu

    LLM pruning has emerged as a promising technology for compressing LLMs, enabling their deployment on resource-limited devices. However, current methodologies typically require access to public calibration samples, which can be challenging to obtain in privacy-sensitive domains. To address this issue, we introduce FedPrLLM, a comprehensive federated pruning f

  47. Tonglong Wei, Yan Lin, Zeyu Zhou, Haomin Wen

    Vehicle GPS trajectories provide valuable movement information that supports various downstream tasks and applications. A desirable trajectory learning model should be able to transfer across regions and tasks without retraining, avoiding the need to maintain multiple specialized models and subpar performance with limited training data. However, each region

  48. Lihong Chen, Hossein Hassani, Soodeh Nikan

    Vision-Language Models (VLMs) have shown remarkable potential in advancing autonomous driving by leveraging multi-modal fusion in order to enhance scene perception, reasoning, and decision-making. Despite their potential, existing models suffer from computational overhead and inefficient integration of multi-view sensor data that make them impractical for re

  49. Abhinaba Roy, Geeta Puri, Dorien Herremans

    We present Text2midi-InferAlign, a novel technique for improving symbolic music generation at inference time. Our method leverages text-to-audio alignment and music structural alignment rewards during inference to encourage the generated music to be consistent with the input caption. Specifically, we introduce two objectives scores: a text-audio consistency

  50. Hanzhuo Tan, Xiaolong Tian, Hanrui Qi, Jiaming Liu

    Recent advances in LLM-based decompilers have been shown effective to convert low-level binaries into human-readable source code. However, there still lacks a comprehensive benchmark that provides large-scale binary-source function pairs, which is critical for advancing the LLM decompilation technology. Creating accurate binary-source mappings incurs severe

  51. Zihan Su, Xuerui Qiu, Hongbin Xu, Tangyu Jiang

    The explosive growth of generative video models has amplified the demand for reliable copyright preservation of AI-generated content. Despite its popularity in image synthesis, invisible generative watermarking remains largely underexplored in video generation. To address this gap, we propose Safe-Sora, the first framework to embed graphical watermarks direc

  52. Huimin Xu, Houjiang Liu, Yan Leng, Ying Ding

    CSCW has long examined how emerging technologies reshape the ways researchers collaborate and produce knowledge, with scientific knowledge production as a central area of focus. As AI becomes increasingly integrated into scientific research, understanding how researchers adapt to it reveals timely opportunities for CSCW research -- particularly in supporting

  53. Ke Chen, Yufei Zhou, Xitong Zhang, Haohan Wang

    Automatic prompt generation plays a crucial role in enabling general-purpose multi-agent systems to perform diverse tasks autonomously. Existing methods typically evaluate prompts based on their immediate task performance, overlooking the intrinsic qualities that determine their reliability. This outcome-centric view not only limits interpretability but also

  54. Ryan Spears, Moonyoung Lee, George Kantor, Oliver Kroemer

    Contact-rich manipulation tasks in agriculture, such as pruning and harvesting, require robots to physically interact with tree structures to maneuver through cluttered foliage. Identifying whether the robot is contacting rigid or soft materials is critical for the downstream manipulation policy to be safe, yet vision alone is often insufficient due to occlu

  55. Ziqing Xing, Zhaoyang Zhang, Zirui Chen, Hongning Ruan

    In this paper, we incorporate physical knowledge into learning-based high-precision target sensing using the multi-view channel state information (CSI) between multiple base stations (BSs) and user equipment (UEs). Such kind of multi-view sensing problem can be naturally cast into a conditional generation framework. To this end, we design a bipartite neural

  56. Xukai Liu, Ye Liu, Shiwen Wu, Yanghai Zhang

    Recent advances in large language models (LLMs) have led to impressive progress in natural language generation, yet their tendency to produce hallucinated or unsubstantiated content remains a critical concern. To improve factual reliability, Retrieval-Augmented Generation (RAG) integrates external knowledge during inference. However, existing RAG systems fac

  57. Tanmay Vilas Samak, Chinmay Vilas Samak, Giovanni Martino, Pranav Nair

    Verification and validation (V&V) of autonomous vehicles (AVs) typically requires exhaustive testing across a variety of operating environments and driving scenarios including rare, extreme, or hazardous situations that might be difficult or impossible to capture in reality. Additionally, physical V&V methods such as track-based evaluations or public-road te

  58. Ziqi Wen, Jonathan Skaza, Shravan Murlidaran, William Y. Wang

    Although models exist that predict human response times (RTs) in tasks such as target search and visual discrimination, the development of image-computable predictors for scene understanding time remains an open challenge. Recent advances in vision-language models (VLMs), which can generate scene descriptions for arbitrary images, combined with the availabil

  59. Jessica Foo, Pradyumna Shyama Prasad, Shaun Khoo

    While the capabilities of large language models (LLMs) have progressed significantly, their use in high-stakes applications have been limited due to risks of hallucination. One key approach in reducing hallucination is retrieval-augmented generation (RAG), but even in such setups, LLMs may still hallucinate when presented with questions outside of the knowle

  60. Xianzhe Dong, Tongxuan Liu, Yuting Zeng, Liangyu Liu

    Multimodal Large Language Models (MLLMs) have been rapidly advancing, enabling cross-modal understanding and generation, and propelling artificial intelligence towards artificial general intelligence. However, existing MLLM inference systems are typically designed based on the architecture of language models, integrating image processing and language process

  61. Shuang Gao, Peter E. Caines

    Transmission Neural Networks (TransNNs) introduced by Gao and Caines (2022) connect virus spread models over networks and neural networks with tuneable activation functions. This paper presents the approximation technique and the underlying assumptions employed by TransNNs in relation to the corresponding Markovian Susceptible-Infected-Susceptible (SIS) mode

  62. Yongchang Gao, Meiling Jin, Zhaofei Yu, Tiejun Huang

    Spike cameras offer unique sensing capabilities but their sparse, asynchronous output challenges semantic understanding, especially for Spike Video-Language Alignment (Spike-VLA) where models like CLIP underperform due to modality mismatch. We introduce SPKLIP, the first architecture specifically for Spike-VLA. SPKLIP employs a hierarchical spike feature ext

  63. Yisheng Zhong, Yizhu Wen, Junfeng Guo, Mehran Kafai

    The protection of cyber Intellectual Property (IP) such as web content is an increasingly critical concern. The rise of large language models (LLMs) with online retrieval capabilities enables convenient access to information but often undermines the rights of original content creators. As users increasingly rely on LLM-generated responses, they gradually dim

  64. Yuxin Lin, Yinglin Zheng, Ming Zeng, Wangzheng Shi

    This paper addresses the gap in predicting turn-taking and backchannel actions in human-machine conversations using multi-modal signals (linguistic, acoustic, and visual). To overcome the limitation of existing datasets, we propose an automatic data collection pipeline that allows us to collect and annotate over 210 hours of human conversation videos. From t

  65. Chihye Han, Michael F. Bonner

    How do different brains create unique visual experiences from identical sensory input? While neural representations vary across individuals, the fundamental architecture underlying these differences remains poorly understood. Here, we reveal that individual visual experience emerges from a high-dimensional neural geometry across the visual cortical hierarchy

  66. Konstantin Batygin, Fred C. Adams

    The formation and early evolution of Jupiter played a pivotal role in sculpting the large-scale architecture of the solar system, intertwining the narrative of Jovian early years with the broader story of the solar system's origins. The details and chronology of Jupiter's formation, however, remain elusive, primarily due to the inherent uncertainties of accr

  67. Sayontan Ghosh, Mahnaz Koupaee, Yash Kumar Lal, Pegah Alipoormolabashi

    Understanding multiparty conversations demands robust Theory of Mind (ToM) capabilities, including the ability to track dynamic information, manage knowledge asymmetries, and distinguish relevant information across extended exchanges. To advance ToM evaluation in such settings, we present a carefully designed scalable methodology for generating high-quality

  68. Kristin Qi, Jiali Cheng, Youxiang Zhu, Hadi Amiri

    Detecting Mild Cognitive Impairment from picture descriptions is critical yet challenging, especially in multilingual and multiple picture settings. Prior work has primarily focused on English speakers describing a single picture (e.g., the 'Cookie Theft'). The TAUKDIAL-2024 challenge expands this scope by introducing multilingual speakers and multiple pictu

  69. Karthik Urs, Jessica Carlson, Aditya Srinivas Manohar, Michael Rakowiecki

    Robotic models are useful for independently varying specific features, but most quadrupedal robots differ so greatly from animal morphologies that they have minimal biomechanical relevance. Commercially available quadrupedal robots are also prohibitively expensive for biological research programs and difficult to customize. Here, we present a low-cost quadru

  70. Tengfei Liu, Haoyang Zhong, Jiazheng Hu, Tan Zhang

    This study presents a dynamic safety margin-based reinforcement learning framework for local motion planning in dynamic and uncertain environments. The proposed planner integrates real-time trajectory optimization with adaptive gap analysis, enabling effective feasibility assessment under partial observability constraints. To address safety-critical computat

  71. Jung Hoon Lee, Sujith Vijayan

    Deep learning (DL) is a powerful tool that can solve complex problems, and thus, it seems natural to assume that DL can be used to enhance the security of wireless communication. However, deploying DL models to edge devices in wireless networks is challenging, as they require significant amounts of computing and power resources. Notably, Spiking Neural Netwo

  72. Tianju Xue

    Differentiable programming is revolutionizing computational science by enabling automatic differentiation (AD) of numerical simulations. While first-order gradients are well-established, second-order derivatives (Hessians) for implicit functions in finite-element-based differentiable physics remain underexplored. This work bridges this gap by deriving and im

  73. Nicola Bogo, Zeyi Zhang, Martin Head-Gordon, Christopher J. Stein

    Charge-transfer excited states are highly relevant for applications in molecular electronics. However, the accurate calculation of these states in large systems is challenging since wave function methods are prohibitively expensive, time-dependent density functional theory with typical functionals is not precise, and the complicated topology of the electroni

  74. Bo Yang, Hengwei Zhang, Jindong Wang, Yuchen Ren

    In surrogate ensemble attacks, using more surrogate models yields higher transferability but lower resource efficiency. This practical trade-off between transferability and efficiency has largely limited existing attacks despite many pre-trained models are easily accessible online. In this paper, we argue that such a trade-off is caused by an unnecessary com

  75. Seonghyeon Moon, Young Sul Cho

    This work extends the thermodynamic analysis of random bond percolation to explosive and hybrid percolation models. We show that this thermodynamic analysis is well applicable to both explosive and hybrid percolation models by using the critical exponents $\alpha$ and $\delta$ obtained from scaling relations with previously measured values of $\beta$ and $\g

  76. Jung Hoon Lee, Sujith Vijayan

    Deep learning (DL) can automatically construct intelligent agents, deep neural networks (alternatively, DL models), that can outperform humans in certain tasks. However, the operating principles of DL remain poorly understood, making its decisions incomprehensible. As a result, it poses a great risk to deploy DL in high-stakes domains in which mistakes or er

  77. Yue Huang, Tianle Hu, Yu Chen, Zi'ang Li

    Single image reflection separation aims to separate the transmission and reflection layers from a mixed image. Existing methods typically combine general priors from pre-trained models with task-specific priors such as text prompts and reflection detection. However, the transmission prior, as the most direct task-specific prior for the target transmission la

  78. Tharaka Wijesundara, Mathew Warren, Nalin Arachchilage

    With the rapid increase in privacy violations in modern software development, regulatory frameworks such as the General Data Protection Regulation (GDPR) have been established to enforce strict data protection practices. However, insufficient privacy awareness among SME software developers contributes to failure in GDPR compliance. For instance, a developer

  79. Hiroshi Yamaguchi, Yudai Miyai, Yuki. Tsubota, Masashi Atira

    The origin of electron-boson interactions is central to understanding high-$T_c$ superconductivity in cuprates. While phonons and magnetic fluctuations are widely considered as candidates for mediating electron pairing, the role of charge fluctuations -- one of the fundamental electronic degrees of freedom -- remains unclear. Here, we investigate the electro

  80. Yifeng Jiao, Yuchen Liu, Yu Zhang, Xin Guo

    The advent of single-cell Assay for Transposase-Accessible Chromatin using sequencing (scATAC-seq) offers an innovative perspective for deciphering regulatory mechanisms by assembling a vast repository of single-cell chromatin accessibility data. While foundation models have achieved significant success in single-cell transcriptomics, there is currently no f

  81. Nan Gao, Xue-Song Lu, Pu Zhang

    In this article we try to recall Claus Michael Ringel's works on the Gorenstein-projective modules. This will involve but not limited to his fundamental contributions, such as in, the solution to the independence problem of totally reflexivity conditions; the technique of $\mho$-quivers; a fast algorithm to obtain the Gorenstein-projective modules over the N

  82. Jiakuan Xie, Pengfei Cao, Yubo Chen, Kang Liu

    Knowledge editing, which aims to update the knowledge encoded in language models, can be deceptive. Despite the fact that many existing knowledge editing algorithms achieve near-perfect performance on conventional metrics, the models edited by them are still prone to generating original knowledge. This paper introduces the concept of "superficial editing" to

  83. Mingqi Shao, Feng Xiong, Zhaoxu Sun, Mu Xu

    Recently, significant advances have been made in 3D object generation. Building upon the generated geometry, current pipelines typically employ image diffusion models to generate multi-view RGB images, followed by UV texture reconstruction through texture baking. While 3D geometry generation has improved significantly, supported by multiple open-source frame

  84. Jiazhu Li, Jian Kuang, Xiaoji Niu

    To overcome the limitation of existing indoor odometry technologies which often cannot simultaneously meet requirements for accuracy cost-effectiveness, and robustness-this paper proposes a novel magnetometer array-aided inertial odometry approach, MSCEKF-MIO (Multi-State Constraint Extended Kalman Filter-based Magnetic-Inertial Odometry). We construct a mag

  85. Alfredo Deaño, Kenneth T-R McLaughlin, Leslie Molag, Nick Simm

    We carry out the asymptotic analysis as $n \to \infty$ of a class of orthogonal polynomials $p_{n}(z)$ of degree $n$, defined with respect to the planar measure \begin{equation*} d\mu(z) = (1-|z|^{2})^{\alpha-1}|z-x|^{\gamma}\mathbf{1}_{|z| < 1}d^{2}z, \end{equation*} where $d^{2}z$ is the two dimensional area measure, $\alpha$ is a parameter that can grow w

  86. Yunseok Jang, Yeda Song, Sungryull Sohn, Lajanugen Logeswaran

    Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have sparked significant interest in developing GUI visual agents. We introduce MONDAY (Mobile OS Navigation Task Dataset for Agents from YouTube), a large-scale dataset of 313K annotated frames from 20K instructional videos capturing diverse real-world mobile OS navigation

  87. Li Lin

    The 3D human pose is vital for modern computer vision and computer graphics, and its prediction has drawn attention in recent years. 3D human pose prediction aims at forecasting a human's future motion from the previous sequence. Ignoring that the arbitrariness of human motion sequences has a firm origin in transition in both temporal and spatial axes limits

  88. Xiangpeng Tian, Xiangyu Liao, Xiao Liu, Meng Li

    All-in-one image restoration aims to recover clear images from various degradation types and levels with a unified model. Nonetheless, the significant variations among degradation types present challenges for training a universal model, often resulting in task interference, where the gradient update directions of different tasks may diverge due to shared par

  89. Yuchang Sun, Yanxi Chen, Yaliang Li, Bolin Ding

    Augmenting large language models (LLMs) with auxiliary tokens has emerged as a promising strategy for enhancing model performance. In this work, we introduce a lightweight method termed latent tokens; these are dummy tokens that may be non-interpretable in natural language but steer the autoregressive decoding process of a Transformer-based LLM via the atten

  90. Beck LaBash, Shahriar Khushrushahi, Fabian Ruehle

    We propose a two-stage deep learning framework for the inverse design of rectangular patch antennas. Our approach leverages generative modeling to learn a latent representation of antenna frequency response curves and conditions a subsequent generative model on these responses to produce feasible antenna geometries. We further demonstrate that leveraging sea

  91. Wanfu Gao, Zengyao Man, Hanlin Pan, Kunpeng Liu

    Feature generation involves creating new features from raw data to capture complex relationships among the original features, improving model robustness and machine learning performance. Current methods using reinforcement learning for feature generation have made feature exploration more flexible and efficient. However, several challenges remain: first, dur

  92. Xuan Wu, Di Wang, Chunguo Wu, Lijie Wen

    Recent studies exploited Large Language Models (LLMs) to autonomously generate heuristics for solving Combinatorial Optimization Problems (COPs), by prompting LLMs to first provide search directions and then derive heuristics accordingly. However, the absence of task-specific knowledge in prompts often leads LLMs to provide unspecific search directions, obst

  93. Ping Xu, Zhiyuan Ning, Pengjiang Li, Wenhao Liu

    Single-cell RNA sequencing (scRNA-seq) reveals cell heterogeneity, with cell clustering playing a key role in identifying cell types and marker genes. Recent advances, especially graph neural networks (GNNs)-based methods, have significantly improved clustering performance. However, the analysis of scRNA-seq data remains challenging due to noise, sparsity, a

  94. Ali Naseh, Harsh Chaudhari, Jaechul Roh, Mingshi Wu

    DeepSeek recently released R1, a high-performing large language model (LLM) optimized for reasoning tasks. Despite its efficient training pipeline, R1 achieves competitive performance, even surpassing leading reasoning models like OpenAI's o1 on several benchmarks. However, emerging reports suggest that R1 refuses to answer certain prompts related to politic

  95. Hansoul Kim, Dong-Ho Lee, Dukyoo Kong, Dong-Soo Kwon

    Robotic endoscopic systems provide intuitive control and eliminate radiation exposure, making them a promising alternative to conventional methods. However, the lack of axial force measurement from the robot remains a major challenge, as it can lead to excessive colonic elongation, perforation, or ureteral complications. Although various methods have been pr

  96. Keisuke Okumura, Hiroki Nagai

    PIBT is a computationally lightweight algorithm that can be applied to a variety of multi-agent pathfinding (MAPF) problems, generating the next collision-free locations of agents given another. Because of its simplicity and scalability, it is becoming a popular underlying scheme for recent large-scale MAPF methods involving several hundreds or thousands of

  97. Tongrui Li, Zhanfeng Liu, Peng Li, Yuzhe Wang

    The interplay between magnetism and electronic band structure is a central theme in condensed matter physics. CeSb, with its complex devil's staircase antiferromagnetic transition, offers a unique opportunity to explore this interplay. Using angle-resolved photoemission spectroscopy (ARPES), we investigate the electronic structure evolution across the devil'

  98. Keqi Deng, Philip C. Woodland

    While Transformer self-attention offers strong parallelism, the Key-Value (KV) cache grows linearly with sequence length and becomes a bottleneck for inference efficiency. Multi-head latent attention was recently developed to compress the KV cache into a low-rank latent space. This paper proposes Multi-head Temporal Latent Attention (MTLA), which further red

  99. João Eduardo Batista, Emil Vatai, Mohamed Wahib

    Large Language Models (LLMs) are increasingly applied in various science domains, yet their broader adoption remains constrained by a critical challenge: the lack of trustworthy, verifiable outputs. Current LLMs often generate answers without reliable source attribution, or worse, with incorrect attributions, posing a barrier to their use in scientific and h

  100. Forrest Mozer, Oleksiy Agapitov, Stuart Bale, John Bonnell

    On November 6, 2024, the Parker Solar Probe flew past Venus to make the first accurate electric field measurement in the nightside Venusian magnetosphere. To achieve this result, the electric field antennas were current biased in a way never before experienced by an electric field detector. This biasing requirement, that the positive bias current in the Venu