March 2025 arXiv papers — page 177
Showing 17,601–17,700 of 23,633 papers
Gilad Abiri
Generative AI is frequently portrayed as revolutionary or even apocalyptic, prompting calls for novel regulatory approaches. This essay argues that such views are misguided. Instead, generative AI should be understood as an evolutionary step in the broader algorithmic media landscape, alongside search engines and social media. Like these platforms, generativ
SGA-INTERACT: A 3D Skeleton-based Benchmark for Group Activity Understanding in Modern Basketball Tactic
cs.CVYuchen Yang, Wei Wang, Yifei Liu, Linfeng Dong
Group Activity Understanding is predominantly studied as Group Activity Recognition (GAR) task. However, existing GAR benchmarks suffer from coarse-grained activity vocabularies and the only data form in single-view, which hinder the evaluation of state-of-the-art algorithms. To address these limitations, we introduce SGA-INTERACT, the first 3D skeleton-base
Linling Kuang, Jiachen Sun, Jin Zhang, Huanxi Cui
As one of the most promising hotspots in the 6G era, space remote sensing information networks play a key and irreplaceable role in areas such as emergency response and scientific research, and are expected to foster remote sensing data processing into the next generation of killer applications. However, due to the inability to deploy ground communication st
Feng Zhang, Yanbin Liu, Weihua Li, Jie Lv
Large Vision and Language Models have exhibited remarkable human-like intelligence in tasks such as natural language comprehension, problem-solving, logical reasoning, and knowledge retrieval. However, training and serving these models require substantial computational resources, posing a significant barrier to their widespread application and further resear
Shinnosuke Matsuo, Riku Togashi, Ryoma Bise, Seiichi Uchida
Active learning (AL) is a label-efficient machine learning paradigm that focuses on selectively annotating high-value instances to maximize learning efficiency. Its effectiveness can be further enhanced by incorporating weak supervision, which uses rough yet cost-effective annotations instead of exact (i.e., full) but expensive annotations. We introduce a no
Abdominal Undulation with Compliant Mechanism Improves Flight Performance of Biomimetic Robotic Butterfly
cs.ROXuyi Lian, Mingyu Luo, Te Lin, Chen Qian
Abdominal Undulation with Compliant Mechanism Improves Flight Performance of Biomimetic Robotic ButterflThis paper presents the design, modeling, and experimental validation of a biomimetic robotic butterfly (BRB) that integrates a compliant mechanism to achieve coupled wing-abdomen motion. Drawing inspiration from the natural f light dynamics of butterflies
Jing Zhang, Zhikai Li, Chengzhi Hu, Xuewen Liu
Segment Anything Model (SAM) exhibits remarkable zero-shot segmentation capability; however, its prohibitive computational costs make edge deployment challenging. Although post-training quantization (PTQ) offers a promising compression solution, existing methods yield unsatisfactory results when applied to SAM, owing to its specialized model components and p
GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow Networks
cs.CLHaoqiang Kang, Enna Sachdeva, Piyush Gupta, Sangjae Bae
Vision-Language Models (VLMs) have recently shown promising advancements in sequential decision-making tasks through task-specific fine-tuning. However, common fine-tuning methods, such as Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) techniques like Proximal Policy Optimization (PPO), present notable limitations: SFT assumes Independent and I
Full Polarization Control of Photons with Evanescent Wave Coupling in the Ultra Subwavelength Gap of Photonic Molecules
physics.opticsRui Zhu, Chenjiang Qian, Shan Xiao, Jingnan Yang
Polarization of photons plays a key role in quantum optics and light-matter interactions, however, it is difficult to control in nanosystems since the eigenstate of a nanophotonic cavity is usually fixed and linearly polarized. Here we reveal polarization control of photons using photonic molecules (PMs) that host supermodes of two coupled nanobeam cavities.
LLM-Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery
cond-mat.mtrl-sciZhilong Song, Qionghua Zhou, Chunjin Ren, Chongyi Ling
Distilling underlying principles from data has historically driven scientific breakthroughs. However, conventional data-driven machine learning often produces complex models that lack interpretability and generalization due to insufficient domain expertise. Here, we present LLM-Feynman, a novel framework that leverages large language models (LLMs) alongside
HFedCKD: Toward Robust Heterogeneous Federated Learning via Data-free Knowledge Distillation and Two-way Contrast
cs.LGYiting Zheng, Bohan Lin, Jinqian Chen, Jihua Zhu
Most current federated learning frameworks are modeled as static processes, ignoring the dynamic characteristics of the learning system. Under the limited communication budget of the central server, the flexible model architecture of a large number of clients participating in knowledge transfer requires a lower participation rate, active clients have uneven
Zhenlong Dai, Bingrui Chen, Zhuoluo Zhao, Xiu Tang
Automated Program Repair (APR) is a task to automatically generate patches for the buggy code. However, most research focuses on generating correct patches while ignoring the consistency between the fixed code and the original buggy code. How to conduct adaptive bug fixing and generate patches with minimal modifications have seldom been investigated. To brid
Robust Optimization Approach for Solving Uncertain Multiobjective Optimization Problems Using the Projected Gradient Method
math.OCShubham Kumar, Nihar Kumar Mahatoa, Debdas Ghosh
Numerous real-world applications of uncertain multiobjective optimization problems (UMOPs) can be found in science, engineering, business, and management. To handle the solution of uncertain optimization problems, robust optimization is a relatively new field. An extended version of the projected gradient method (PGM) for a deterministic smooth multiobjectiv
Quanjian Song, Zhihang Lin, Zhanpeng Zeng, Ziyue Zhang
Existing camera motion-controlled video generation methods face computational bottlenecks in fine-tuning and inference. This paper proposes LightMotion, a light and tuning-free method for simulating camera motion in video generation. Operating in the latent space, it eliminates additional fine-tuning, inpainting, and depth estimation, making it more streamli
Lei Liu, Xiujuan Zhang, Ming-Hui Lu, Yan-Feng Chen
Skyrmions--topologically protected nanoscale spin textures with vortex-like configurations--hold transformative potential for ultra-dense data storage, spintronics and quantum computing. However, their practical utility is challenged by dynamic instability, complex interaction, and the lack of deterministic control. While recent efforts using classical wave
Timo Aukusti Laine
Large Language Models (LLMs) encode semantic relationships in high-dimensional vector embeddings. This paper explores the analogy between LLM embedding spaces and quantum mechanics, positing that LLMs operate within a quantized semantic space where words and phrases behave as quantum states. To capture nuanced semantic interference effects, we extend the sta
Amir Mohammad Izadi, Seyed Mohammad Hadi Hosseini, Soroush Vafaie Tabar, Ali Abdollahi
Text-to-image generative models have made significant advancements in recent years; however, accurately capturing intricate details in textual prompts-such as entity missing, attribute binding errors, and incorrect relationships remains a formidable challenge. In response, we present an innovative, training-free method that directly addresses these challenge
Romain Lacombe
Evolution-based protein structure prediction models have achieved breakthrough success in recent years. However, they struggle to generalize beyond evolutionary priors and on sequences lacking rich homologous data. Here we present a novel, out-of-domain benchmark based on sactipeptides, a rare class of ribosomally synthesized and post-translationally modifie
T. Zalialiutdinov, D. Solovyev
In this study, we develop and implement a specialized coupled-cluster (CC) approach tailored for accurately describing atoms and molecules in strong magnetic fields. Using the open-source Ghent Quantum Chemistry Package (\texttt{GQCP}) in conjunction with the Python-based Simulations of Chemistry Framework (\texttt{PySCF}), we calculate potential energy curv
Xirui Hu, Jiahao Wang, Hao Chen, Weizhan Zhang
Recent advances in text-to-image generation have driven interest in generating personalized human images that depict specific identities from reference images. Although existing methods achieve high-fidelity identity preservation, they are generally limited to single-ID scenarios and offer insufficient facial editability. We present DynamicID, a tuning-free
Dynamics of Light Localization via Coherent Control: The Interplay of Transmission, Absorption and Disorder in Photonic Crystals
physics.opticsNancy Ghangas, Ghanasyam Remesh, Venu Gopal Achanta, Shubhrangshu Dasgupta
This study investigates the interplay between structural disorder, absorption, and Lyapunov exponent dynamics to exploit localization phenomena in photonic crystals with engineered defect layers. We generate disorder by introducing random refractive index variations in one of the bilayers, while the application of a control field to $\Lambda$-type atoms with
Constraints on the Scale Parameter of Regular Black Hole in Asymptotically Safe Gravity from Extreme Mass Ratio Inspirals
gr-qcLai Zhao, Meirong Tang, Zhaoyi Xu
This paper evaluates the potential for constraining the quantum scale parameter $\xi$ of regular black hole within the asymptotically safe gravity framework using gravitational waves from extreme mass ratio inspirals (EMRIs). Since $\xi$ cannot be precisely determined from first principles, observational constraints become crucial. We employ the Augmented An
Xiaofeng Xue
In this paper, we prove a fluctuation theorem for the occupation time of the multi-species stirring process on a lattice starting from a stationary distribution. Our result shows that the occupation times of different species interact with each other at the level of equilibrium fluctuation. The proof of our result utilizes the resolvent strategy introduced i
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
cs.CVHuaqi Tao, Bingxi Liu, Calvin Chen, Tingjun Huang
Visual Place Recognition (VPR) is a crucial capability for long-term autonomous robots, enabling them to identify previously visited locations using visual information. However, existing methods remain limited in indoor settings due to the highly repetitive structures inherent in such environments. We observe that scene texts frequently appear in indoor spac
Yanwei Huang, Yan Miao, Di Weng, Adam Perer
Data profiling plays a critical role in understanding the structure of complex datasets and supporting numerous downstream tasks, such as social media analytics and financial fraud detection. While existing research predominantly focuses on structured data formats, a substantial portion of semi-structured textual data still requires ad-hoc and arduous manual
Xukun Zhou, Fengxin Li, Ming Chen, Yan Zhou
Audio-driven human gesture synthesis is a crucial task with broad applications in virtual avatars, human-computer interaction, and creative content generation. Despite notable progress, existing methods often produce gestures that are coarse, lack expressiveness, and fail to fully align with audio semantics. To address these challenges, we propose ExGes, a n
Haixiang zhang, Mengyu Cao, Mei Lu, Jiaying Song
Let $V$ be an $n$-dimensional vector space over the finite field $\mathbb{F}_q$ and ${V\brack k}$ denote the family of all $k$-dimensional subspaces of $V$. A family $\mathcal{F}\subseteq {V\brack k}$ is called $k$-uniform $r$-wise $t$-intersecting if for any $F_1, F_2, \dots, F_r \in \mathcal{F}$, we have $\dim\left(\bigcap_{i=1}^r F_i \right) \geq t$. An $
Enming Zhang, Peizhe Gong, Xingyuan Dai, Min Huang
Ensuring the safety of vision-language models (VLMs) in autonomous driving systems is of paramount importance, yet existing research has largely focused on conventional benchmarks rather than safety-critical evaluation. In this work, we present SCD-Bench (Safety Cognition Driving Benchmark) a novel framework specifically designed to assess the safety cogniti
Wei-Jie Zhang, Zhenyu Zhang, Jifeng Hu, Bing-Nan Lu
Finite-volume extrapolation is an important step for extracting physical observables from lattice calculations. However, it is a significant challenge for the system with long-range interactions. We employ symbolic regression to regress finite-volume extrapolation formula for both short-range and long-range interactions. The regressed formula still holds the
Alsharif Abuadbba, Sean Lamont, Ejaz Ahmed, Cody Christopher
As malware detection evolves, attackers adopt sophisticated evasion tactics. Traditional file-level fingerprinting, such as cryptographic and fuzzy hashes, is often overlooked as a target for evasion. Malware variants exploit minor binary modifications to bypass detection, as seen in Microsoft's discovery of GoldMax variations (2020-2021). However, no large-
Imon Kalyan, Ieng Wai Un, Gilles Rosolen, Nir Shitrit
For decades, there have been multiple seemingly contradicting experimental reports on the dependence of the photoluminescence from metal nanostructures on their size. We reconcile these reports using a simple analytic formula which is found to match well photoluminescence measurements for a range of structures and illumination conditions. Our expression requ
Mushfiqur Rahman, Ismail Guvenc, David Ramirez, Chau-Wai Wong
Deployment of cellular networks in urban areas requires addressing various challenges. For example, high-rise buildings with varying geometrical shapes and heights contribute to signal attenuation, reflection, diffraction, and scattering effects. This creates a high possibility of coverage holes (CHs) within the proximity of the buildings. Detecting these CH
Sirinda Palahan
The rise of online programming education has necessitated more effective, personalized interactions, a gap that PythonPal aims to fill through its innovative learning system integrated with a chatbot. This research delves into PythonPal's potential to enhance the online learning experience, especially in contexts with high student-to-teacher ratios where the
Relationships between Students' Social Roles and Academic Performance based on Social Network Analysis
cs.SISirinda Palahan
Peer interaction and social roles have been important factors in students' academic performance. Recent work on what influences academic performance in Thailand has focused on the quality of a school, students' backgrounds, and students themselves. A few works have analyzed the correlation between social roles and students' academic achievement. Therefore, t
VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
cs.CLYanling Wang, Yihan Zhao, Xiaodong Chen, Shasha Guo
Large vision-language models (LVLMs) have demonstrated remarkable achievements, yet the generation of non-factual responses remains prevalent in fact-seeking question answering (QA). Current multimodal fact-seeking benchmarks primarily focus on comparing model outputs to ground truth answers, providing limited insights into the performance of modality-specif
Jean Seo, Jaeyoon Kim, Hyopil Shin
We propose the Mixture of Frozen Experts (MoFE) architecture, which integrates Parameter-efficient Fine-tuning (PEFT) and the Mixture of Experts (MoE) architecture to enhance both training efficiency and model scalability. By freezing the Feed Forward Network (FFN) layers within the MoE framework, MoFE significantly reduces the number of trainable parameters
Mingwei Wang, Yingtian Liu, Junheng Peng, Yong Li
Seismic data acquisition is often affected by various types of noise, which degrade data quality and hinder subsequent interpretation. Recovery of seismic data becomes particularly challenging in the presence of strong noise, which significantly impacts both data accuracy and geological analysis. This study proposes a novel single-encoder, multiple-decoder n
Improving Access to Trade and Investment Information in Thailand through Intelligent Document Retrieval
cs.IRSirinda Palahan
Overseas investment and trade can be daunting for beginners due to the vast amount of complex information. This paper presents a chatbot system that integrates natural language processing and information retrieval techniques to simplify the document retrieval process. The proposed system identifies the most relevant content, enabling users to navigate the in
Shijun Cheng, Mohammad H. Taufik, Tariq Alkhalifah
Current neural operators often struggle to generalize to complex, out-of-distribution conditions, limiting their ability in seismic wavefield representation. To address this, we propose a generative neural operator (GNO) that leverages generative diffusion models (GDMs) to learn the underlying statistical distribution of scattered wavefields while incorporat
A Study of Effectiveness of Brand Domain Identification Features for Phishing Detection in 2025
cs.CRRina Mishra, Gaurav Varshney
Phishing websites continue to pose a significant security challenge, making the development of robust detection mechanisms essential. Brand Domain Identification (BDI) serves as a crucial step in many phishing detection approaches. This study systematically evaluates the effectiveness of features employed over the past decade for BDI, focusing on their weigh
Muhammad Ahmed Mohsin, Ahsan Bilal, Sagnik Bhattacharya, John M. Cioffi
Future wireless networks aim to deliver high data rates and lower power consumption while ensuring seamless connectivity, necessitating robust optimization. Large language models (LLMs) have been deployed for generalized optimization scenarios. To take advantage of generative AI (GAI) models, we propose retrieval augmented generation (RAG) for multi-sensor w
Varad Vaidya, Jishnu Keshavan
Due to dynamic variations such as changing payload, aerodynamic disturbances, and varying platforms, a robust solution for quadrotor trajectory tracking remains challenging. To address these challenges, we present a deep reinforcement learning (DRL) framework that achieves physical dynamics invariance by directly optimizing force/torque inputs, eliminating t
Cong Chen, Mingyu Liu, Chenchen Jing, Yizhou Zhou
This paper aims to address the challenge of hallucinations in Multimodal Large Language Models (MLLMs) particularly for dense image captioning tasks. To tackle the challenge, we identify the current lack of a metric that finely measures the caption quality in concept level. We hereby introduce HalFscore, a novel metric built upon the language graph and is de
A Mesh Is Worth 512 Numbers: Spectral-domain Diffusion Modeling for High-dimension Shape Generation
cs.CVJiajie Fan, Amal Trigui, Andrea Bonfanti, Felix Dietrich
Recent advancements in learning latent codes derived from high-dimensional shapes have demonstrated impressive outcomes in 3D generative modeling. Traditionally, these approaches employ a trained autoencoder to acquire a continuous implicit representation of source shapes, which can be computationally expensive. This paper introduces a novel framework, spect
Xiao Wang, Yuehang Li, Fuling Wang, Bo Jiang
Accurate sign language understanding serves as a crucial communication channel for individuals with disabilities. Current sign language translation algorithms predominantly rely on RGB frames, which may be limited by fixed frame rates, variable lighting conditions, and motion blur caused by rapid hand movements. Inspired by the recent successful application
Density-Matrix Embedding Based Multi-Configurational Perturbation Theory Approach to Single-Ion Magnets
physics.chem-phZhe-Bin Guan, Hong Jiang
Multi-configurational wave-function theory (MC-WFT) that combines complete active space self-consistent field (CASSCF) approach with subsequent state interaction (SI) treatment of spin-orbit coupling (SOC), abbreviated as CASSCF-SO, plays important roles in microscopic understanding of single-ion magnets (SIMs) with different central transition metal or lant
PathVQ: Reforming Computational Pathology Foundation Model for Whole Slide Image Analysis via Vector Quantization
cs.CVHonglin Li, Zhongyi Shui, Yunlong Zhang, Chenglu Zhu
Computational pathology and whole-slide image (WSI) analysis are pivotal in cancer diagnosis and prognosis. However, the ultra-high resolution of WSIs presents significant modeling challenges. Recent advancements in pathology foundation models have improved performance, yet most approaches rely on [CLS] token representation of tile ViT as slide-level inputs
Assessment of the point-wise approach for the Turbulent Settling of finite-size particles
physics.flu-dynFrancesco Battistaa, Sergio Chibbarob, Paolo Gualtieria
We study the settling of suspensions of relatively large particles with a diameter of the order of ten Kolmogorov scales and density slightly larger than the carrier fluid in statistically steady homogeneous isotropic turbulence. The particle-to-fluid density ratio is varied to obtain a wide range of Galileo numbers, which are the ratios between buoyancy and
Pankaj Popli, Ananyo Maitra, Sriram Ramaswamy
We present a detailed analytical and numerical examination, on square and triangular lattices, of the non-reciprocal planar spin model introduced in Dadhichi et al., Phys. Rev. E 101, 052601 (2020). We show that the effect of lattice anisotropy should persist at large scales, leading to a ``mass'' for the angle field of the spins, and behaviour not in the ``
ExKG-LLM: Leveraging Large Language Models for Automated Expansion of Cognitive Neuroscience Knowledge Graphs
cs.AIAli Sarabadani, Kheirolah Rahsepar Fard, Hamid Dalvand
The paper introduces ExKG-LLM, a framework designed to automate the expansion of cognitive neuroscience knowledge graphs (CNKG) using large language models (LLMs). It addresses limitations in existing tools by enhancing accuracy, completeness, and usefulness in CNKG. The framework leverages a large dataset of scientific papers and clinical reports, applying
Pengyu Yang, Xin Zhang, Song Lin
To address the issue of excessive quantum resource requirements in Kuperberg's algorithm for the dihedral hidden subgroup problem, this paper proposes a distributed algorithm based on the function decomposition. By splitting the original function into multiple subfunctions and distributing them to multiple quantum nodes for parallel processing, the algorithm
Chuheng Wei, Ziye Qin, Siyan Li, Ziyan Zhang
Driving behavior is inherently personal, influenced by individual habits, decision-making styles, and physiological states. However, most existing datasets treat all drivers as homogeneous, overlooking driver-specific variability. To address this gap, we introduce the Personalized Driving Behavior (PDB) dataset, a multi-modal dataset designed to capture pers
Global Convergence and Rate Analysis of the Steepest Descent Method for Uncertain Multiobjective Optimization via a Robust Optimization Approach
math.OCShubham Kumar, Nihar Kumar Mahato, Debdas Ghosh
In this article, we extend our previous work (Applicable Analysis, 2024, pp. 1-25) on the steepest descent method for uncertain multiobjective optimization problems. While that study established local convergence, it did not address global convergence and the rate of convergence of the steepest descent algorithm. To bridge this gap, we provide rigorous proof
Accodemy: AI Powered Code Learning Platform to Assist Novice Programmers in Overcoming the Fear of Coding
cs.HCM. A. F. Aamina, V. Kavishcan, W. M. P. B. B. Jayaratne, K. K. D. S. N. Kannangara
Computer programming represents a rapidly evolving and sought-after career path in the 21st century. Nevertheless, novice learners may find the process intimidating for several reasons, such as limited and highly competitive career opportunities, peer and parental pressure for academic success, and course difficulties. These factors frequently contribute to
SKG-LLM: Developing a Mathematical Model for Stroke Knowledge Graph Construction Using Large Language Models
cs.AIAli Sarabadani, Kheirolah Rahsepar Fard, Hamid Dalvand
The purpose of this study is to introduce SKG-LLM. A knowledge graph (KG) is constructed from stroke-related articles using mathematical and large language models (LLMs). SKG-LLM extracts and organizes complex relationships from the biomedical literature, using it to increase the accuracy and depth of KG in stroke research. In the proposed method, GPT-4 was
Zhefan Wang, Huanjun Kong, Jie Ying, Wanli Ouyang
Large language models (LLMs) commonly struggle with specialized or emerging topics which are rarely seen in the training corpus. Graph-based retrieval-augmented generation (GraphRAG) addresses this by structuring domain knowledge as a graph for dynamic retrieval. However, existing pipelines involve complex engineering workflows, making it difficult to isolat
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
cs.CVYuxuan Luo, Jiaqi Tang, Chenyi Huang, Feiyang Hao
Chinese calligraphy, a UNESCO Heritage, remains computationally challenging due to visual ambiguity and cultural complexity. Existing AI systems fail to contextualize their intricate scripts, because of limited annotated data and poor visual-semantic alignment. We propose CalliReader, a vision-language model (VLM) that solves the Chinese Calligraphy Contextu
Qiaole Dong, Yanwei Fu
Dense point tracking is a challenging task requiring the continuous tracking of every point in the initial frame throughout a substantial portion of a video, even in the presence of occlusions. Traditional methods use optical flow models to directly estimate long-range motion, but they often suffer from appearance drifting without considering temporal consis
Optimal Transport for Brain-Image Alignment: Unveiling Redundancy and Synergy in Neural Information Processing
q-bio.NCYang Xiao, Wang Lu, Jie Ji, Ruimeng Ye
The design of artificial neural networks (ANNs) is inspired by the structure of the human brain, and in turn, ANNs offer a potential means to interpret and understand brain signals. Existing methods primarily align brain signals with stimulus signals using Mean Squared Error (MSE), which focuses only on local point-wise alignment and ignores global matching,
Fei Tang, Yongliang Shen, Hang Zhang, Siqi Chen
Humans can flexibly switch between different modes of thinking based on task complexity: from rapid intuitive judgments to in-depth analytical understanding. However, current Graphical User Interface (GUI) grounding systems which locate interface elements based on natural language instructions rely solely on immediate prediction without reasoning, struggling
George Tang, Aditya Agarwal, Weiqiao Han, Trevor Darrell
We generalize lifting to semantic lifting by incorporating per-view masks that indicate relevant pixels for lifting tasks. These masks are determined by querying corresponding multiscale pixel-aligned feature maps, which are derived from scene representations such as distilled feature fields and feature point clouds. However, storing per-view feature maps re
Mobility-Aware Multi-Task Decentralized Federated Learning for Vehicular Networks: Modeling, Analysis, and Optimization
cs.NIDongyu Chen, Tao Deng, He Huang, Juncheng Jia
Federated learning (FL) is a promising paradigm that can enable collaborative model training between vehicles while protecting data privacy, thereby significantly improving the performance of intelligent transportation systems (ITSs). In vehicular networks, due to mobility, resource constraints, and the concurrent execution of multiple training tasks, how to
SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic Prompts
cs.CVShijia Zhao, Qiming Xia, Xusheng Guo, Pufan Zou
Recently, sparsely-supervised 3D object detection has gained great attention, achieving performance close to fully-supervised 3D objectors while requiring only a few annotated instances. Nevertheless, these methods suffer challenges when accurate labels are extremely absent. In this paper, we propose a boosting strategy, termed SP3D, explicitly utilizing the
L. C. Eze, R. Jajcay, T. Jajcayová, D. Závacká
The Cage Problem requires for a given pair $k \geq 3, g \geq 3$ of integers the determination of the order of a smallest $k$-regular graph of girth $g$. We address a more general version of this problem and look for the $(k,g)$-spectrum of orders of $(k,g)$-graphs: the (infinite) list of all orders of $(k,g)$-graphs. By establishing these spectra we aim to g
Qiqi Lin, Xiaoyang Ji, Shengfang Zhai, Qingni Shen
Large language models (LLMs) have achieved remarkable success in natural language processing, yet their performance and computational costs vary significantly. LLM routers play a crucial role in dynamically balancing these trade-offs. While previous studies have primarily focused on routing efficiency, security vulnerabilities throughout the entire LLM route
Proper Characterization of Heat-to-Electric Conversion Efficiency of Liquid Thermogalvanic Cells
physics.app-phQiangqiang Huang, Yuchi Chen, Ronggui Yang, Xin Qian
Liquid thermogalvanic cells (LTCs) have emerged as a promising technology for harvesting low-grade heat due to their low cost, compact design, and high thermopower. However, discrepancies exist in quantifying their output power and efficiency. The commonly used figure of merit, ZT = S^2\sigma T/k, is based on electrolyte properties but fails to account for e
Detecting Correlation Efficiently in Stochastic Block Models: Breaking Otter's Threshold in the Entire Supercritical Regime
cs.DSGuanyi Chen, Jian Ding, Shuyang Gong, Zhangsong Li
Consider a pair of sparse correlated stochastic block models $\mathcal S(n,\tfrac{\lambda}{n},\epsilon;s)$ subsampled from a common parent stochastic block model with two symmetric communities, average degree $\lambda=O(1)$, divergence parameter $\epsilon\in (0,1)$ and subsampling probability $s$. For all $\epsilon\in(0,1)$ and $\Delta>0$, we construct a sta
AXAI-CDSS : An Affective Explainable AI-Driven Clinical Decision Support System for Cannabis Use
cs.HCTongze Zhang, Tammy Chung, Anind Dey, Sang Won Bae
As cannabis use has increased in recent years, researchers have come to rely on sophisticated machine learning models to predict cannabis use behavior and its impact on health. However, many artificial intelligence (AI) models lack transparency and interpretability due to their opaque nature, limiting their trust and adoption in real-world medical applicatio
StructGS: Adaptive Spherical Harmonics and Rendering Enhancements for Superior 3D Gaussian Splatting
cs.CVZexu Huang, Min Xu, Stuart Perry
Recent advancements in 3D reconstruction coupled with neural rendering techniques have greatly improved the creation of photo-realistic 3D scenes, influencing both academic research and industry applications. The technique of 3D Gaussian Splatting and its variants incorporate the strengths of both primitive-based and volumetric representations, achieving sup
Seungju Cho, Hongsin Lee, Changick Kim
Adversarial training significantly enhances adversarial robustness, yet superior performance is predominantly achieved on balanced datasets. Addressing adversarial robustness in the context of unbalanced or long-tailed distributions is considerably more challenging, mainly due to the scarcity of tail data instances. Previous research on adversarial robustnes
Maximal coin-position entanglement and non-Hermitian skin effect in discrete-time quantum walks
quant-phDing Cheng, Yi Li, Hao Zhao, Haijun Kang
A distinctive feature of non-Hermitian systems is the skin effect, which has attracted widespread attention in recent studies. Quantum walks provide a powerful platform for exploring the underlying mechanisms of the non-Hermitian skin effect. Additionally, the generation of hybrid entanglement in quantum walks is recognized as another crucial property. Howev
Hariharan Narayanan, Piyush Srivastava
Polynomial-time deterministic approximation of volumes of polytopes, up to an approximation factor that grows at most sub-exponentially with the dimension, remains an open problem. Recent work on this question has focused on identifying interesting classes of polytopes for which such approximation algorithms can be obtained. In this paper, we focus on one su
Guanyu Cao, Takuya Maekawa, Kazuya Ohara, Yasue Kishino
This study proposes a new deep learning method for reconstructing depth images of moving objects within a specific area using Wi-Fi channel state information (CSI). The Wi-Fi-based depth imaging technique has novel applications in domains such as security and elder care. However, reconstructing depth images from CSI is challenging because learning the mappin
Yanbiao Ma, Wei Dai, Wenke Huang, Jiayi Chen
Data heterogeneity in federated learning, characterized by a significant misalignment between local and global distributions, leads to divergent local optimization directions and hinders global model training. Existing studies mainly focus on optimizing local updates or global aggregation, but these indirect approaches demonstrate instability when handling h
Chengxuan Qian, Kai Han, Jiaxin Liu, Zhenlong Yuan
Multimodal learning integrates complementary information from diverse modalities to enhance the decision-making process. However, the potential of multimodal collaboration remains under-exploited due to disparities in data quality and modality representation capabilities. To address this, we introduce DynCIM, a novel dynamic curriculum learning framework des
Yunfeng Li, Xiaolin Li Zhitao Li, Gangqiang Li
With the booming development of prosumers, there is an urgent need for a prosumer energy management system to take full advantage of the flexibility of prosumers and take into account the interests of other parties. However, building such a system will undoubtedly reveal users' privacy. In this paper, by solving the non-independent and identical distribution
Yihong Xu, Quan Zhou
The challenges posed by high-dimensional data and use of the simplex constraint are two major concerns in the empirical application of the synthetic control method (SCM) in econometric studies. To address both issues simultaneously, we propose a Bayesian SCM that integrates a soft simplex constraint within spike-and-slab variable selection. The hierarchical
Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation Variation
cs.CRShengfang Zhai, Jiajun Li, Yue Liu, Huanran Chen
In recent years, text-to-image (T2I) diffusion models have gained significant attention for their ability to generate high quality images reflecting text prompts. However, their growing popularity has also led to the emergence of backdoor threats, posing substantial risks. Currently, effective defense strategies against such threats are lacking due to the di
Ruibin Xu, Zheng Zheng, Yanying Liang, Zhu-Jun Zheng
The advancement of classical machine learning is inherently linked to the establishment and progression of classical dataset. In quantum machine learning (QML), there is an analogous imperative for the development of quantum entangled datasets comprised with huge quantity and high quality. Especially for multipartite mixed-state datasets, due to the lack of
A Quantitative Evaluation of the Expressivity of BMI, Pose and Gender in Body Embeddings for Recognition and Identification
cs.CVBasudha Pal, Siyuan Huang, Rama Chellappa
Person Re-identification (ReID) systems that match individuals across images or video frames are essential in many real-world applications. However, existing methods are often influenced by attributes such as gender, pose, and body mass index (BMI), which vary in unconstrained settings and raise concerns related to fairness and generalization. To address thi
Jun Tamura, Yuki Itaya, Kenichi Hayashi, Kouji Yamamoto
Classification problems are essential statistical tasks that form the foundation of decision-making across various fields, including patient prognosis and treatment strategies for critical conditions. Consequently, evaluating the performance of classification models is of significant importance, and numerous evaluation metrics have been proposed. Among these
Xiaoyu Cheng, Yangyang Hu, Tianci Miao, Wenbo Liu
The complementary field-effect transistors (CFETs), featuring vertically stacked n/p-FETs, enhance integration density and significantly reduce the area of standard cells such as static random-access memory (SRAM). However, the advantage of area scaling through CFETs is hindered by the imbalance in N/P transistor counts (typically 4N/2P) within SRAM cells. I
Zi Ye, Kai Yu, Song Lin
Graph Convolutional Networks (GCNs) are specialized neural networks for feature extraction from graph-structured data. In contrast to traditional convolutional networks, GCNs offer distinct advantages when processing irregular data, which is ubiquitous in real-world applications. This paper introduces an enhancement to GCNs based on spectral methods by integ
Mingxiang Cao, Weiying Xie, Xin Zhang, Jiaqing Zhang
Multi-modal fusion holds great promise for integrating information from different modalities. However, due to a lack of consideration for modal consistency, existing multi-modal fusion methods in the field of remote sensing still face challenges of incomplete semantic information and low computational efficiency in their fusion designs. Inspired by the obser
Zuqing Li, Junhao Gan, Jianzhong Qi
Diffusion-based tabular data synthesis models have yielded promising results. However, when the data dimensionality increases, existing models tend to degenerate and may perform even worse than simpler, non-diffusion-based models. This is because limited training samples in high-dimensional space often hinder generative models from capturing the distribution
Mobility-Aware Decentralized Federated Learning with Joint Optimization of Local Iteration and Leader Selection for Vehicular Networks
cs.NIDongyu Chen, Tao Deng, Juncheng Jia, Siwei Feng
Federated learning (FL) emerges as a promising approach to empower vehicular networks, composed by intelligent connected vehicles equipped with advanced sensing, computing, and communication capabilities. While previous studies have explored the application of FL in vehicular networks, they have largely overlooked the intricate challenges arising from the mo
Yu Liu, Hao Tang, Haiqi Zhang, Jing Qin
Out-of-distribution (OOD) detection is crucial for ensuring the reliability and safety of machine learning models in real-world applications. While zero-shot OOD detection, which requires no training on in-distribution (ID) data, has become feasible with the emergence of vision-language models like CLIP, existing methods primarily focus on semantic matching
Identifying Evidence Subgraphs for Financial Risk Detection via Graph Counterfactual and Factual Reasoning
cs.CEHuaming Du, Lei Yuan, Qing Yang, Xingyan Chen
Company financial risks pose a significant threat to personal wealth and national economic stability, stimulating increasing attention towards the development of efficient andtimely methods for monitoring them. Current approaches tend to use graph neural networks (GNNs) to model the momentum spillover effect of risks. However, due to the black-box nature of
Yu Wang, Qingmei Zhao
The global null controllability of stochastic semilinear parabolic equations with globally Lipschitz nonlinearities has been addressed in recent literature. However, there are no results concerning their numerical approximation and the behavior of discrete controls when the discretization parameter goes to zero. This paper is intended to studying the null co
Generalizable Machine Learning Models for Predicting Data Center Server Power, Efficiency, and Throughput
cs.LGNuoa Lei, Arman Shehabi, Jun Lu, Zhi Cao
In the rapidly evolving digital era, comprehending the intricate dynamics influencing server power consumption, efficiency, and performance is crucial for sustainable data center operations. However, existing models lack the ability to provide a detailed and reliable understanding of these intricate relationships. This study employs a machine learning-based
ELT-METIS imaging simulations for disks and envelopes associated with FU Ori-type objects
astro-ph.IMMichihiro Takami, Gilles Otten, Olivier Absil, Christian Delacroix
We investigate the detectability of extended mid-infrared (MIR) emission associated with FU-Ori type objects (FUors) using the METIS coronagraphs on the 39-m Extremely Large Telescope (ELT). The imaging simulations were made for three representative filters ($\lambda$=3.8, 4.8, and 11.3 micron) of the METIS instrument. We demonstrate that the detectability o
Physics-Informed Residual Neural Ordinary Differential Equations for Enhanced Tropical Cyclone Intensity Forecasting
physics.ao-phFan Meng
Accurate tropical cyclone (TC) intensity prediction is crucial for mitigating storm hazards, yet its complex dynamics pose challenges to traditional methods. Here, we introduce a Physics-Informed Residual Neural Ordinary Differential Equation (PIR-NODE) model to precisely forecast TC intensity evolution. This model leverages the powerful non-linear fitting c
OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection
cs.CVAdrian Chow, Evelien Riddell, Yimu Wang, Sean Sedwards
Open-vocabulary 3D object detection for autonomous driving aims to detect novel objects beyond the predefined training label sets in point cloud scenes. Existing approaches achieve this by connecting traditional 3D object detectors with vision-language models (VLMs) to regress 3D bounding boxes for novel objects and perform open-vocabulary classification thr
Yi Zhang, Zhao Pan
Droplet-fiber interactions, prevalent in nature and widely applied across various engineering fields, have garnered significant research interest. Many works have focused on the interactions between droplets and single or two fibers. However, the wetting behavior of droplets, especially the maximum droplets that can be retained, on fiber hubs formed by many
Qidong Su, Wei Zhao, Xin Li, Muralidhar Andoorveedu
To improve the efficiency of distributed large language model (LLM) inference, various parallelization strategies, such as tensor and pipeline parallelism, have been proposed. However, the distinct computational characteristics inherent in the two stages of LLM inference-prefilling and decoding-render a single static parallelization strategy insufficient for
Xiaoyu Chen, HongSheng Hu
We prove that there exists a bound $N'_L(W)$ for a positively weighted Coxeter group $(W, S, L)$ of finite rank. In particular, Lusztig's $\boldsymbol{a}$-function of $(W, S, L)$ is bounded.
Mingrui Zhang, Xiaowu Dai, Lexin Li
The kidney paired donation (KPD) program provides an innovative solution to overcome incompatibility challenges in kidney transplants by matching incompatible donor-patient pairs and facilitating kidney exchanges. To address unequal access to transplant opportunities, there are two widely used fairness criteria: group fairness and individual fairness. Howeve
Zhangchi Qiu, Linhao Luo, Zicheng Zhao, Shirui Pan
Conversational Recommender Systems (CRSs) have emerged as a transformative paradigm for offering personalized recommendations through natural language dialogue. However, they face challenges with knowledge sparsity, as users often provide brief, incomplete preference statements. While recent methods have integrated external knowledge sources to mitigate this
Junwei Zhang, Xuewu Chang, Ping Jin, Lei Wang
In 2007, J. P. Cossey conjectured that if $G$ is a finite $p$-solvable group and $\varphi$ is an irreducible Brauer character of $G$ with vertex $Q$, then the number of lifts of $\varphi$ is at most $|Q:Q'|$. In this paper we revisited Cossey's conjecture for $p=2$ from the perspective of Navarro vertices and obtained a new way to count the number of lifts o
Tianshu Huang, Arjun Ramesh, Emily Ruppel, Nuno Pereira
Accurately estimating workload runtime is a longstanding goal in computer systems, and plays a key role in efficient resource provisioning, latency minimization, and various other system management tasks. Runtime prediction is particularly important for managing increasingly complex distributed systems in which more sophisticated processing is pushed to the