May 2024 arXiv papers — page 35
Showing 3,401–3,500 of 20,894 papers
Lang Feng, Pengjie Gu, Bo An, Gang Pan
Diffusion planners have shown promise in handling long-horizon and sparse-reward tasks due to the non-autoregressive plan generation. However, their inherent stochastic risk of generating infeasible trajectories presents significant challenges to their reliability and stability. We introduce a novel approach, the Trajectory Aggregation Tree (TAT), to address
Dongjae Jeon, Wonje Jeung, Taeheon Kim, Albert No
Machine unlearning (MU) aims to remove the influence of specific data from trained models, addressing privacy concerns and ensuring compliance with regulations such as the ``right to be forgotten.'' Evaluating strong unlearning, where the unlearned model is indistinguishable from one retrained without the forgetting data, remains a significant challenge in d
Mingyuan Liu, Lu Xu, Shengnan Liu, Jicong Zhang
The success of Large Vision Models (LVMs) is accompanied by vast data volumes, which are prohibitively expensive in medical diagnosis.To address this, recent efforts exploit Parameter-Efficient Fine-Tuning (PEFT), which trains a small number of weights while freezing the rest.However, they typically assign trainable weights to the same positions in LVMs in a
SelMatch: Effectively Scaling Up Dataset Distillation via Selection-Based Initialization and Partial Updates by Trajectory Matching
cs.CVYongmin Lee, Hye Won Chung
Dataset distillation aims to synthesize a small number of images per class (IPC) from a large dataset to approximate full dataset training with minimal performance loss. While effective in very small IPC ranges, many distillation methods become less effective, even underperforming random sample selection, as IPC increases. Our examination of state-of-the-art
Yingqi Liu, Yifan Shi, Qinglun Li, Baoyuan Wu
Personalized Federated Learning (PFL) is proposed to find the greatest personalized models for each client. To avoid the central failure and communication bottleneck in the server-based FL, we concentrate on the Decentralized Personalized Federated Learning (DPFL) that performs distributed model training in a Peer-to-Peer (P2P) manner. Most personalized work
BO4IO: A Bayesian optimization approach to inverse optimization with uncertainty quantification
math.OCYen-An Lu, Wei-Shou Hu, Joel A. Paulson, Qi Zhang
This work addresses data-driven inverse optimization (IO), where the goal is to estimate unknown parameters in an optimization model from observed decisions that can be assumed to be optimal or near-optimal solutions to the optimization problem. The IO problem is commonly formulated as a large-scale bilevel program that is notoriously difficult to solve. Dev
D. van der Sluis
To investigate whether "Intelligence is the capacity of an information-processing system to adapt to its environment while operating with insufficient knowledge and resources", we look at utilising the non axiomatic reasoning system (NARS) for speech recognition. This article presents NUTS: raNdom dimensionality redUction non axiomaTic reasoning few Shot lea
MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization
cs.CVTianchen Zhao, Xuefei Ning, Tongcheng Fang, Enshu Liu
Diffusion models have achieved significant visual generation quality. However, their significant computational and memory costs pose challenge for their application on resource-constrained mobile devices or even desktop GPUs. Recent few-step diffusion models reduces the inference time by reducing the denoising steps. However, their memory consumptions are st
HFGS: 4D Gaussian Splatting with Emphasis on Spatial and Temporal High-Frequency Components for Endoscopic Scene Reconstruction
cs.CVHaoyu Zhao, Xingyue Zhao, Lingting Zhu, Weixi Zheng
Robot-assisted minimally invasive surgery benefits from enhancing dynamic scene reconstruction, as it improves surgical outcomes. While Neural Radiance Fields (NeRF) have been effective in scene reconstruction, their slow inference speeds and lengthy training durations limit their applicability. To overcome these limitations, 3D Gaussian Splatting (3D-GS) ba
Xin Xiao, Bohong Wu, Jiacong Wang, Chunyuan Li
Existing image-text modality alignment in Vision Language Models (VLMs) treats each text token equally in an autoregressive manner. Despite being simple and effective, this method results in sub-optimal cross-modal alignment by over-emphasizing the text tokens that are less correlated with or even contradictory with the input images. In this paper, we advoca
Enda Yu, Dezun Dong, Xiangke Liao
In distributed deep learning, communication remains a critical bottleneck. While modern hardware advances rapidly, over 60 percent of production HPC systems still rely on legacy infrastructure (V100 GPUs, multi-plane Ethernet/InfiniBand), necessitating communication optimization without hardware upgrades. Existing approaches face three key limitations: (1) s
Vadim Voevodkin, Andrey Sokolov
Modern storage systems intensively utilize data prefetching algorithms while processing sequences of the read requests. Performance of the prefetching algorithm (for instance increase of the cache hit ratio of the cache system - CHR) directly affects overall performance characteristics of the storage system (read latency, IOPS, etc.). There are widely known
Nonreciprocal singularities dominated by the dissipative photon-magnon coupling in non-Hermitian systems
cond-mat.mes-hallYongzhang Shi, Chi Zhang, Zhenhui Hao, Changjun Jiang
We investigated the magnon-photon coupling in an open cavity magnonic system, which leads to two different nonreciprocal singularities dominated by the dissipative coupling. One type of singularity is the exceptional point, which is just on the exceptional surface in parameter space. The other type of singularity is the bound state in the continuum discovere
An algorithm applied the Turing pattern model to control active swarm robots using only information from neighboring modules
cs.ROTakeshi Ishida
Swarm robots, inspired by the emergence of animal herds, are robots that assemble a large number of modules and self-organize themselves to form specific morphologies and exhibit specific functions. These modular robots perform relatively simple actions and controls, and create macroscopic morphologies and functions through the interaction of a large number
Adjustable Robust Nonlinear Network Design Without Controllable Elements under Load Scenario Uncertainties
math.OCJohannes Thürauf, Julia Grübel, Martin Schmidt
We study network design problems for nonlinear and nonconvex flow models without controllable elements under load scenario uncertainties, i.e., under uncertain injections and withdrawals. To this end, we apply the concept of adjustable robust optimization to compute a network design that admits a feasible transport for all, possibly infinitely many, load sce
Geetha Ramasubbu, André Kaup, Christian Herglotz
The Bj{\o}ntegaard Delta rate (BD-rate) objectively assesses the coding efficiency of video codecs using the rate-distortion (R-D) performance but overlooks encoding energy, which is crucial in practical applications, especially for those on handheld devices. Although R-D analysis can be extended to incorporate encoding energy as energy-distortion (E-D), it
Andrii Liashyk, Nicolai Reshetikhin, Ivan Sechin
In this paper, we develop the framework for quantum integrable systems on an integrable classical background. We call them hybrid quantum integrable systems (hybrid integrable systems), and we show that they occur naturally in the semiclassical limit of quantum integrable systems. We start with an outline of the concept of hybrid dynamical systems. Then we g
Ferromagnetic ferroelectricity due to the Kugel-Khomskii mechanism of the orbital ordering assisted by atomic Hund's second rule effects
cond-mat.str-elI. V. Solovyev, R. Ono, S. A. Nikolaev
The exchange interactions in insulators depend on the orbital state of magnetic ions, obeying certain phenomenological principles, known as Goodenough-Kanamori-Anderson rules. Particularly, the ferro order of alike orbitals tends to stabilize antiferromagnetic interactions, while the antiferro order of unlike orbitals favors ferromagnetic interactions. The K
Sohan Ghodla
Gravitational radiation alone is not efficient in hardening the orbit of a wide binary black hole (BBH). By employing a toy model for the interstellar medium (ISM) surrounding BBHs, here we discuss the effect of this baryonic medium on BBH dynamics. Depending on the BBH's mass, we show that a binary surrounded by an isotropic cold neutral medium (i.e., an as
Towards robust prediction of material properties for nuclear reactor design under scarce data -- a study in creep rupture property
cs.LGYu Chen, Edoardo Patelli, Zhen Yang, Adolphus Lye
Advances in Deep Learning bring further investigation into credibility and robustness, especially for safety-critical engineering applications such as the nuclear industry. The key challenges include the availability of data set (often scarce and sparse) and insufficient consideration of the uncertainty in the data, model, and prediction. This paper therefor
Strong limits on keV-scale galactic sterile neutrino dark matter with stray light from NuSTAR after 11 years of operation
hep-phR. A. Krivonos, V. V. Barinov, A. A. Mukhin, D. S. Gorbunov
Using tremendous photon statistics gained with the stray light aperture of the NuSTAR telescope over 11 years of operation, we set strong limits on the emission of close to monochromatic photons from the radiative decays of putative dark matter sterile neutrinos in the Milky Way. In the energy range of 3-20 keV covered by the NuSTAR, the obtained limits reac
JUNO Collaboration, Angel Abusleme, Thomas Adam, Kai Adamowicz
This paper presents an energy resolution study of the JUNO experiment, incorporating the latest knowledge acquired during the detector construction phase. The determination of neutrino mass ordering in JUNO requires an exceptional energy resolution better than 3\% at 1~MeV. To achieve this ambitious goal, significant efforts have been undertaken in the desig
Yangxiao Lu, Jishnu Jaykumar P, Yunhui Guo, Nicholas Ruozzi
Novel Instance Detection and Segmentation (NIDS) aims at detecting and segmenting novel object instances given a few examples of each instance. We propose a unified, simple, yet effective framework (NIDS-Net) comprising object proposal generation, embedding creation for both instance templates and proposal regions, and embedding matching for instance label a
Mohamed Saadi, Felix Kölzow, Christian Kontermann, Matthias Oechsner
We quantify the uncertainty of the L\"ammer model of damage evolution when fitted to (noisy) observations of damage evolution in cyclic fatigue experiments with and without dwell time. We therefore develop a bootstrap method by sampling over blocks of load cycles in the experiments in order to quantify the uncertainty in the material parameters of the L\"amm
Nikolai Terekhov, Maksim Zhukovskii
Given a graph $F$ and a positive integer $n$, the weak $F$-saturation number $\mathrm{wsat}(K_n,F)$ is the minimum number of edges in a graph $H$ on $n$ vertices such that the edges missing in $H$ can be added, one at a time, so that every edge creates a copy of $F$. Kalai in 1985 introduced a linear algebraic approach that became one of the most efficient t
The structure of Ferroelectric BaBiO$_3$/BaTiO$_3$ Interfaces grown by Molecular Beam Epitaxy
cond-mat.mtrl-sciMerve Baksi, Divine P. Kumah
We investigate the lattice structure of heterostructures comprising of ferroelectric BaTiO$_3$ (BTO) thin films and BaBiO$_3$ (BBO), the insulating parent-compound of the high Tc superconductor. Motivated by theoretical predictions of exotic phenomena in BBO-based heterostructures including interfacial conductivity, superconductivity and topologically-protec
Johannes van der Vyver
Fare evasion is a problem for public transport companies, with LSTM models this issue can help companies get an analytical insight into where this issue occurs the most, to prevent capital loss. In addition to the financial burden this problem causes, having more inspectors is not enough to alleviate the problem. The purpose of this study is to find a differ
Takeshi Ikeda, Takafumi Kouno, Yusuke Nakayama, Kohei Yamaguchi
We study Schubert calculus in the torus-equivariant quantum $K$-ring of the Lagrangian Grassmannian $\mathrm{LG}(n)$. Our main tool is the $K$-theoretic Peterson map due to Kato. The map is from the (localized) equivariant $K$-homology ring $K_{*}^{T}(\mathrm{Gr}_{G})$ of the affine Grassmannian $\mathrm{Gr}_{G}$ of the symplectic group $G=\mathrm{Sp}_{2n}(\
White dwarf magnetospheres: Shielding volatile content of icy objects and implications for volatile pollution scarcity
astro-ph.EPWen-Han Zhou, Shang-Fei Liu, Douglas N. C. Lin
Context. About 25% -- 50% of white dwarfs are found to be contaminated by heavy elements, which are believed to originate from external sources such as planetary materials. Elemental abundances suggest that most of the pollutants are rocky objects and only a small fraction of white dwarfs bear traces of volatile accretion. Aims. In order to account for the s
Qizhang Li, Yiwen Guo, Wangmeng Zuo, Hao Chen
Adversarial prompts generated using gradient-based methods exhibit outstanding performance in performing automatic jailbreak attacks against safety-aligned LLMs. Nevertheless, due to the discrete nature of texts, the input gradient of LLMs struggles to precisely reflect the magnitude of loss change that results from token replacements in the prompt, leading
Yin Shi, Xiaomei Zhang, Alexey Arefiev, Baifei Shen
Low-intensity light beams carrying Orbital Angular Momentum (OAM), commonly known as vortex beams, have garnered significant attention due to promising applications in areas ranging from optical trapping to communication. In recent years, there has been a surge in global research exploring the potential of high-intensity vortex laser beams and specifically t
Nelleke Bunji, Bartosz Fornal, Kassandra Garcia
The nature of dark matter remains one of the greatest unsolved mysteries in elementary particle physics. It might well be that the dark matter particle belongs to a dark sector completely secluded or extremely weakly coupled to the visible sector. We demonstrate that gravitational waves arising from first order phase transitions in the early Universe can be
Alina Akhmiarova, Alexander Veretennikov
Three versions of the Weak Law of Large Numbers are proposed for weakly dependent and generally speaking non-equally distributed random variables, with finite or possibly infinite expectations.
Xing Hu, Yuan Cheng, Dawei Yang, Zhihang Yuan
Post-training quantization (PTQ) serves as a potent technique to accelerate the inference of large language models (LLMs). Nonetheless, existing works still necessitate a considerable number of floating-point (FP) operations during inference, including additional quantization and de-quantization, as well as non-linear operators such as RMSNorm and Softmax. T
Eduard N. Tsoy, Laziz A. Suyunov
Our analysis suggests strongly that stationary pulses exist in nonlinear media with second-, third-, and fourth-order dispersion. A theory, based on the variational approach, is developed for finding approximate parameters of such solitons. It is obtained that the soliton velocity in the retarded reference frame can be different from the inverse of the group
Sean P. Svihla, Manuel E. Lladser
Consider a tree $T=(V,E)$ with root $\circ$ and edge length function $\ell:E\to\mathbb{R}_+$. The phylogenetic covariance matrix of $T$ is the matrix $C$ with rows and columns indexed by $L$, the leaf set of $T$, with entries $C(i,j):=\sum_{e\in[i\wedge j,o]}\ell(e)$, for each $i,j\in L$. Recent work [15] has shown that the phylogenetic covariance matrix of
Yong Qi, Gabriel Kyebambo, Siyuan Xie, Wei Shen
Safety limitations in service robotics across various industries have raised significant concerns about the need for robust mechanisms ensuring that robots adhere to safe practices, thereby preventing actions that might harm humans or cause property damage. Despite advances, including the integration of Knowledge Graphs (KGs) with Large Language Models (LLMs
A System for Quantifying Data Science Workflows with Fine-Grained Procedural Logging and a Pilot Study
cs.HCJinjin Zhao, Avidgor Gal, Sanjay Krishnan
It is important for researchers to understand precisely how data scientists turn raw data into insights, including typical programming patterns, workflow, and methodology. This paper contributes a novel system, called DataInquirer, that tracks incremental code executions in Jupyter notebooks (a type of computational notebook). The system allows us to quantit
Exploring Thermography Technology: A Comprehensive Facial Dataset for Face Detection, Recognition, and Emotion
cs.CVMohamed Fawzi Abdelshafie Abuhussein, Ashraf Darwish, Aboul Ella Hassanien
This dataset includes 6823 thermal images captured using a UNI-T UTi165A camera for face detection, recognition, and emotion analysis. It consists of 2485 facial recognition images depicting emotions (happy, sad, angry, natural, surprised), 2054 images for face recognition, and 2284 images for face detection. The dataset covers various conditions, color pale
Enhancing Sliding Performance with Aerial Robots: Analysis and Solutions for Non-Actuated Multi-Wheel Configurations
cs.ROTong Hui, Jefferson Ghielmini, Dimitrios Papageorgiou, Marco Tognon
Sliding tasks performed by aerial robots are valuable for inspection and simple maintenance tasks at height, such as non-destructive testing and painting. Although various end-effector designs have been used for such tasks, non-actuated wheel configurations are more frequently applied thanks to their rolling capability for sliding motion, mechanical simplici
David Zhou, Sarah Sterman
In each step of the creative writing process, writers must grapple with their creative goals and individual perspectives. This process affects the writer's sense of authenticity and their engagement with the written output. Fluent text generation by AIs risks undermining the reflective loop of rewriting. We hypothesize that deliberately generating imperfect
MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
cs.CVAkio Hayakawa, Masato Ishii, Takashi Shibuya, Yuki Mitsufuji
This study aims to construct an audio-video generative model with minimal computational cost by leveraging pre-trained single-modal generative models for audio and video. To achieve this, we propose a novel method that guides single-modal models to cooperatively generate well-aligned samples across modalities. Specifically, given two pre-trained base diffusi
Constrained monotone mean--variance investment-reinsurance under the Cram\'er--Lundberg model with random coefficients
q-fin.PMXiaomin Shi, Zuo Quan Xu
This paper studies an optimal investment-reinsurance problem for an insurer (she) under the Cram\'er--Lundberg model with monotone mean--variance (MMV) criterion. At any time, the insurer can purchase reinsurance (or acquire new business) and invest in a security market consisting of a risk-free asset and multiple risky assets whose excess return rate and vo
Andrew H. Lee, Sina J. Semnani, Galo Castillo-López, Gäel de Chalendar
Creating multilingual task-oriented dialogue (TOD) agents is challenging due to the high cost of training data acquisition. Following the research trend of improving training data efficiency, we show for the first time, that in-context learning is sufficient to tackle multilingual TOD. To handle the challenging dialogue state tracking (DST) subtask, we break
Alka Luqman, Shivanshu Shekhar, Anupam Chattopadhyay
This work integrates peer-to-peer federated learning tools with NS3, a widely used network simulator, to create a novel simulator designed to allow heterogeneous device experiments in federated learning. This cross-platform adaptability addresses a critical gap in existing simulation tools, enhancing the overall utility and user experience. NS3 is leveraged
Enabling Generative Design Tools with LLM Agents for Mechanical Computation Devices: A Case Study
cs.HCQiuyu Lu, Jiawei Fang, Zhihao Yao, Yue Yang
In the field of Human-Computer Interaction (HCI), interactive devices with embedded mechanical computation are gaining attention. The rise of these cutting-edge devices has created a need for specialized design tools that democratize the prototyping process. While current tools streamline prototyping through parametric design and simulation, they often come
Zavareh Bozorgasl, Hao Chen
This paper presents the development and application of Wavelet Kolmogorov-Arnold Networks (Wav-KAN) in federated learning. We implemented Wav-KAN \cite{wav-kan} in the clients. Indeed, we have considered both continuous wavelet transform (CWT) and also discrete wavelet transform (DWT) to enable multiresolution capabaility which helps in heteregeneous data di
Deform3DGS: Flexible Deformation for Fast Surgical Scene Reconstruction with Gaussian Splatting
cs.CVShuojue Yang, Qian Li, Daiyun Shen, Bingchen Gong
Tissue deformation poses a key challenge for accurate surgical scene reconstruction. Despite yielding high reconstruction quality, existing methods suffer from slow rendering speeds and long training times, limiting their intraoperative applicability. Motivated by recent progress in 3D Gaussian Splatting, an emerging technology in real-time 3D rendering, thi
Caio Kalil Lauand, Sean Meyn
Many machine learning and optimization algorithms are built upon the framework of stochastic approximation (SA), for which the selection of step-size (or learning rate) $\{\alpha_n\}$ is crucial for success. An essential condition for convergence is the assumption that $\sum_n \alpha_n = \infty$. Moreover, in all theory to date it is assumed that $\sum_n \al
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
cs.CLYuanpu Cao, Tianrong Zhang, Bochuan Cao, Ziyi Yin
Researchers have been studying approaches to steer the behavior of Large Language Models (LLMs) and build personalized LLMs tailored for various applications. While fine-tuning seems to be a direct solution, it requires substantial computational resources and may significantly affect the utility of the original LLM. Recent endeavors have introduced more ligh
Mike Steel
A wide variety of stochastic models of cladogenesis (based on speciation and extinction) lead to an identical distribution on phylogenetic tree shapes once the edge lengths are ignored. By contrast, the distribution of the tree's edge lengths is generally quite sensitive to the underlying model. In this paper, we review the impact of different model choices
Machine Learning-Driven Optimization of TPMS Architected Materials Using Simulated Annealing
cond-mat.mtrl-sciAkshansh Mishra
The research paper presents a novel approach to optimizing the tensile stress of Triply Periodic Minimal Surface (TPMS) structures through machine learning and Simulated Annealing (SA). The study evaluates the performance of Random Forest, Decision Tree, and XGBoost models in predicting tensile stress, using a dataset generated from finite element analysis o
Tao Wang, Sylvia Herbert, Sicun Gao
Policy gradient methods have enabled deep reinforcement learning (RL) to approach challenging continuous control problems, even when the underlying systems involve highly nonlinear dynamics that generate complex non-smooth optimization landscapes. We develop a rigorous framework for understanding how policy gradient methods mollify non-smooth optimization la
Agarose Derived Carbon Based Nanocomposite for Hydrogen Storage at Near-Ambient Conditions
cond-mat.mtrl-sciA Flamina, R M Raghavendra, Anandh Subramaniam, Raghupathy Yuvaraj
Nanocomposites comprising of high surface area adsorption materials and nanosized transition metals have emerged as a promising strategy for hydrogen storage application due to their inherent ability to store atomic and molecular forms of hydrogen by invoking mechanisms like physisorption and spillover mechanism or Kubas interaction. The potential use of the
Chengyuan Liu, Yangyang Kang, Shihang Wang, Lizhi Qing
The performance on general tasks decreases after Large Language Models (LLMs) are fine-tuned on domain-specific tasks, the phenomenon is known as Catastrophic Forgetting (CF). However, this paper presents a further challenge for real application of domain-specific LLMs beyond CF, called General Capabilities Integration (GCI), which necessitates the integrati
LDMol: A Text-to-Molecule Diffusion Model with Structurally Informative Latent Space Surpasses AR Models
cs.LGJinho Chang, Jong Chul Ye
With the emergence of diffusion models as a frontline generative model, many researchers have proposed molecule generation techniques with conditional diffusion models. However, the unavoidable discreteness of a molecule makes it difficult for a diffusion model to connect raw data with highly complex conditions like natural language. To address this, here we
Yuecheng Zhang, Guanhua Fang, Wen Yu
Clustering of event stream data is of great importance in many application scenarios, including but not limited to, e-commerce, electronic health, online testing, mobile music service, etc. Existing clustering algorithms fail to take outlier data into consideration and are implemented without theoretical guarantees. In this paper, we propose a robust tempora
Yimeng Liu, Misha Sra
Choreography creation requires high proficiency in artistic and technical skills. Choreographers typically go through four stages to create a dance piece: preparation, studio, performance, and reflection. This process is often individualized, complicated, and challenging due to multiple constraints at each stage. To assist choreographers, most prior work has
Robin de Jong, Farbod Shokrieh
We discuss canonical local heights on abelian varieties over non-archimedean fields from the point of view of Berkovich analytic spaces. Our main result is a refinement of N\'eron's classical result relating canonical local heights with intersection multiplicities on the N\'eron model. We also revisit Tate's explicit formulas for N\'eron's canonical local he
Seokil Ham, Sangmin Woo, Jin-Young Kim, Hyojun Go
We present Diffusion Model Patching (DMP), a simple method to boost the performance of pre-trained diffusion models that have already reached convergence, with a negligible increase in parameters. DMP inserts a small, learnable set of prompts into the model's input space while keeping the original model frozen. The effectiveness of DMP is not merely due to t
mTREE: Multi-Level Text-Guided Representation End-to-End Learning for Whole Slide Image Analysis
cs.CVQuan Liu, Ruining Deng, Can Cui, Tianyuan Yao
Multi-modal learning adeptly integrates visual and textual data, but its application to histopathology image and text analysis remains challenging, particularly with large, high-resolution images like gigapixel Whole Slide Images (WSIs). Current methods typically rely on manual region labeling or multi-stage learning to assemble local representations (e.g.,
Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-Action
cs.CLZhenyu Pan, Haozheng Luo, Manling Li, Han Liu
We present a Conversational Chain-of-Action (Conv-CoA) framework for Open-domain Conversational Question Answering (OCQA). Compared with literature, Conv-CoA addresses three major challenges: (i) unfaithful hallucination that is inconsistent with real-time or domain facts, (ii) weak reasoning performance in conversational scenarios, and (iii) unsatisfying pe
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
cs.CVSangmin Woo, Jaehyuk Jang, Donguk Kim, Yubin Choi
Recent advancements in Large Vision Language Models (LVLMs) have revolutionized how machines understand and generate textual responses based on visual inputs, yet they often produce "hallucinatory" outputs that misinterpret visual information, posing challenges in reliability and trustworthiness. We propose RITUAL, a simple decoding method that reduces hallu
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
cs.CVSangmin Woo, Donguk Kim, Jaehyuk Jang, Yubin Choi
Large Vision Language Models (LVLMs) demonstrate strong capabilities in visual understanding and description, yet often suffer from hallucinations, attributing incorrect or misleading features to images. We observe that LVLMs disproportionately focus on a small subset of image tokens--termed blind tokens--which are typically irrelevant to the query (e.g., ba
Arnab Char, T. Karthick
Given a graph $G$, the parameters $\chi(G)$ and $\omega(G)$ respectively denote the chromatic number and the clique number of $G$. A function $f : \mathbb{N} \rightarrow \mathbb{N}$ such that $f(1) = 1$ and $f(x) \geq x$, for all $x \in \mathbb{N}$ is called a $\chi$-binding function for the given class of graphs $\cal{G}$ if every $G \in \cal{G}$ satisfies
Jacy Anthis, Kristian Lum, Michael Ekstrand, Avi Feller
The rise of general-purpose artificial intelligence (AI) systems, particularly large language models (LLMs), has raised pressing moral questions about how to reduce bias and ensure fairness at scale. Researchers have documented a sort of "bias" in the significant correlations between demographics (e.g., race, gender) in LLM prompts and responses, but it rema
Hyperspectral and multispectral image fusion with arbitrary resolution through self-supervised representations
cs.CVTing Wang, Zipei Yan, Jizhou Li, Xile Zhao
The fusion of a low-resolution hyperspectral image (LR-HSI) with a high-resolution multispectral image (HR-MSI) has emerged as an effective technique for achieving HSI super-resolution (SR). Previous studies have mainly concentrated on estimating the posterior distribution of the latent high-resolution hyperspectral image (HR-HSI), leveraging an appropriate
Benchmarking Skeleton-based Motion Encoder Models for Clinical Applications: Estimating Parkinson's Disease Severity in Walking Sequences
cs.CVVida Adeli, Soroush Mehraban, Irene Ballester, Yasamin Zarghami
This study investigates the application of general human motion encoders trained on large-scale human motion datasets for analyzing gait patterns in PD patients. Although these models have learned a wealth of human biomechanical knowledge, their effectiveness in analyzing pathological movements, such as parkinsonian gait, has yet to be fully validated. We pr
Yingwen Wu, Ruiji Yu, Xinwen Cheng, Zhengbao He
In the open world, detecting out-of-distribution (OOD) data, whose labels are disjoint with those of in-distribution (ID) samples, is important for reliable deep neural networks (DNNs). To achieve better detection performance, one type of approach proposes to fine-tune the model with auxiliary OOD datasets to amplify the difference between ID and OOD data th
Haogeng Liu, Quanzeng You, Xiaotian Han, Yongfei Liu
In the realm of Multimodal Large Language Models (MLLMs), vision-language connector plays a crucial role to link the pre-trained vision encoders with Large Language Models (LLMs). Despite its importance, the vision-language connector has been relatively less explored. In this study, we aim to propose a strong vision-language connector that enables MLLMs to a
Akshat Mohan Dasula, Hrushitha Tigulla, Preethika Bhukya
Traditionally in the domain of legal research, the retrieval of pertinent citations from intricate case descriptions has demanded manual effort and keyword-based search applications that mandate expertise in understanding legal jargon. Legal case descriptions hold pivotal information for legal professionals and researchers, necessitating more efficient and a
Hanjun Luo, Ziye Deng, Ruizhe Chen, Zuozhu Liu
The rapid development and reduced barriers to entry for Text-to-Image (T2I) models have raised concerns about the biases in their outputs, but existing research lacks a holistic definition and evaluation framework of biases, limiting the enhancement of debiasing techniques. To address this issue, we introduce FAIntbench, a holistic and precise benchmark for
The Impacts of Data, Ordering, and Intrinsic Dimensionality on Recall in Hierarchical Navigable Small Worlds
cs.IROwen Pendrigh Elliott, Jesse Clark
Vector search systems, pivotal in AI applications, often rely on the Hierarchical Navigable Small Worlds (HNSW) algorithm. However, the behaviour of HNSW under real-world scenarios using vectors generated with deep learning models remains under-explored. Existing Approximate Nearest Neighbours (ANN) benchmarks and research typically has an over-reliance on s
Verónica Becher, Tomás Tropea
Fix a finite alphabet. A necklace is a circular word. For positive integers $n$ and~$k$, a necklace is $(n,k)$-perfect if all words of length $n$ occur $k$ times but at positions with different congruence modulo $k$, for any convention of the starting position. We define the notion of a Lyndon pair and we use it to construct the lexicographically greatest $(
Xiangjun Gao, Xiaoyu Li, Yiyu Zhuang, Qi Zhang
Neural 3D representations such as Neural Radiance Fields (NeRF), excel at producing photo-realistic rendering results but lack the flexibility for manipulation and editing which is crucial for content creation. Previous works have attempted to address this issue by deforming a NeRF in canonical space or manipulating the radiance field based on an explicit me
A new class of evolution multivalued quasi-variational inequalities I: existence and nonsmooth optimal control
math.FAShengda Zeng, Vicenţiu D. Rădulescu
In this paper, we consider a new kind of evolution multivalued quasi-variational inequalities with feedback effect and a nonlinear bifunction which contain several (evolution) quasi-variational/hemivariational inequalities as special cases. The main contribution of this paper is twofold. The first goal is to establish a novel framework for proving the existe
Chenyang Le, Yao Qian, Dongmei Wang, Long Zhou
There is a rising interest and trend in research towards directly translating speech from one language to another, known as end-to-end speech-to-speech translation. However, most end-to-end models struggle to outperform cascade models, i.e., a pipeline framework by concatenating speech recognition, machine translation and text-to-speech models. The primary c
Correlation effects in magic-angle twisted bilayer graphene: An auxiliary-field quantum Monte Carlo study
cond-mat.str-elZhi-Yu Xiao, Shiwei Zhang
Magic angle twisted bilayer graphene (MATBG) presents a fascinating platform for investigating the effects of electron interactions in topological flat bands. The Bistritzer-MacDonald (BM) model provides a simplified quantitative description of the flat bands. Introducing long-range Coulomb interactions leads to an interacting BM (IBM) Hamiltonian, a momentu
Jin-Lei Yang, Hai-Bin Zhang, Tai-Fu Feng
A flavor-dependent model (FDM) is proposed in this work. The model extends the Standard Model by an extra $U(1)_F$ local gauge group, two scalar doublets, one scalar singlet and two right-handed neutrinos, where the additional $U(1)_F$ charges are related to the particles' flavor. The new fermion sector in the FDM can explain the flavor mixings puzzle and th
Zheng Tracy Ke, Jingming Wang
Topic modeling is a widely utilized tool in text analysis. We investigate the optimal rate for estimating a topic model. Specifically, we consider a scenario with $n$ documents, a vocabulary of size $p$, and document lengths at the order $N$. When $N\geq c\cdot p$, referred to as the long-document case, the optimal rate is established in the literature at $\
Souvik Jana, Shasvath J Kapadia, Tejaswi Venumadhav, Surhud More
We present a detailed exposition of a statistical method for estimating cosmological parameters from the observation of a large number of strongly lensed binary-black-hole (BBH) mergers observable by next (third) generation (XG) gravitational-wave (GW) detectors. This method, first presented in Jana (2023 Phys. Rev. Lett. 130 261401), compares the observed n
Wei Li, Houfeng Wang
Grammatical error correction (GEC) is a task dedicated to rectifying texts with minimal edits, which can be decoupled into two components: detection and correction. However, previous works have predominantly focused on direct correction, with no prior efforts to integrate both into a single model. Moreover, the exploration of the detection-correction paradig
Searching for the highest energy of pulsation and critical luminosity of Swift J0243.6+6124 observed by Insight-HXMT
astro-ph.HEQing-Xia Zhao, Xian Hou, Ming-Yu Ge, Shuang-Nan Zhang
Owing to the broad energy coverage of Insight-HXMT in the hard X-ray band, we detected the highest energy of pulsation exceeding 200 keV around the 2017-2018 outburst peak of the first Galactic pulsating ultraluminous X-ray source (PULX) Swift J0243.6+6124, which is the highest energy detected from PULXs to date. We also obtained the highest energy of pulsat
Yuanle Mo, Xin Hong, Bowen Gao, Yinjun Jia
Protein-protein interactions are central mediators in many biological processes. Accurately predicting the effects of mutations on interactions is crucial for guiding the modulation of these interactions, thereby playing a significant role in therapeutic development and drug discovery. Mutations generally affect interactions hierarchically across three level
Itamar Cohen
Caching is extensively used in various networking environments to optimize performance by reducing latency, bandwidth, and energy consumption. To optimize performance, caches often advertise their content using indicators, which are data structures that trade space efficiency for accuracy. However, this tradeoff introduces the risk of false indications. Exis
Alia Hamieh, Habiba Kadiri, Nathan Ng, Greg Martin
This is an ongoing list of problems that has resulted from the PIMS (Pacific Institute of Mathematical Sciences) Collaborative Research Group L-functions in Analytic Number Theory: 2022- 2025. The focus of this list is on Moments of $L$-functions and related topics.
Yudong Wang, Damai Dai, Zhifang Sui
Most work treats large language models as black boxes without in-depth understanding of their internal working mechanism. In order to explain the internal representations of LLMs, we propose a gradient-based metric to assess the activation level of model parameters. Based on this metric, we obtain three preliminary findings. (1) When the inputs are in the sa
Fumian Chen, Hui Fang
Ranking algorithms as an essential component of retrieval systems have been constantly improved in previous studies, especially regarding relevance-based utilities. In recent years, more and more research attempts have been proposed regarding fairness in rankings due to increasing concerns about potential discrimination and the issue of echo chamber. These a
Dania Mezher, Moussa Daamouch
Seymour Second Neighborhood Conjecture (SSNC) asserts that every finite oriented graph has a vertex whose second out-neighborhood is at least as large as its first out-neighborhood. Such a vertex is called a Seymour vertex. A digraph $D = (V, E)$ is $k$-anti-transitive if for every pair of vertices $u, v \in V$, the existence of a directed path of length $k$
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
cs.LGJian Qian, Haichen Hu, David Simchi-Levi
Motivated by the recent discovery of a statistical and computational reduction from contextual bandits to offline regression (Simchi-Levi and Xu, 2021), we address the general (stochastic) Contextual Markov Decision Process (CMDP) problem with horizon H (as known as CMDP with H layers). In this paper, we introduce a reduction from CMDPs to offline density es
Mingjia Yin, Hao Wang, Wei Guo, Yong Liu
The sequential recommender (SR) system is a crucial component of modern recommender systems, as it aims to capture the evolving preferences of users. Significant efforts have been made to enhance the capabilities of SR systems. These methods typically follow the model-centric paradigm, which involves developing effective models based on fixed datasets. Howev
LNS2+RL: Combining Multi-Agent Reinforcement Learning with Large Neighborhood Search in Multi-Agent Path Finding
cs.ROYutong Wang, Tanishq Duhan, Jiaoyang Li, Guillaume Sartoretti
Multi-Agent Path Finding (MAPF) is a critical component of logistics and warehouse management, which focuses on planning collision-free paths for a team of robots in a known environment. Recent work introduced a novel MAPF approach, LNS2, which proposed to repair a quickly obtained set of infeasible paths via iterative replanning, by relying on a fast, yet l
Yongjae Lee, Zhaoliang Zhang, Deliang Fan
3D Gaussian Splatting (3DGS) has made significant strides in novel view synthesis. However, its suboptimal densification process results in the excessively large number of Gaussian primitives, which impacts frame-per-second and increases memory usage, making it unsuitable for low-end devices. To address this issue, many follow-up studies have proposed variou
JUNO Collaboration, Angel Abusleme, Thomas Adam, Kai Adamowicz
We explore the decay of bound neutrons into invisible particles (e.g., $n\rightarrow 3 \nu$ or $nn \rightarrow 2 \nu$) in the JUNO liquid scintillator detector, which do not produce an observable signal. The invisible decay includes two decay modes: $ n \rightarrow { inv} $ and $ nn \rightarrow { inv} $. The invisible decays of $s$-shell neutrons in $^{12}{\
Classical and quantum thermodynamics in a non-equilibrium regime: Application to Stirling engine
cond-mat.stat-mechShoki Koyanagi, Yoshitaka Tanimura
We have developed a thermodynamic theory in the non-equilibrium regime, which we describe as a thermodynamic system-bath model [S. Koyanagi and Y. Tanimura, J. Chem. Phys. \textbf{160}, 234112 (2024)]. Based on the dimensionless (DL) minimum work principle, non-equilibrium thermodynamic potentials are expressed in terms of non-equilibrium extensive and inten
Weizhen He, Yiheng Deng, Yunfeng Yan, Feng Zhu
Human intelligence can retrieve any person according to both visual and language descriptions. However, the current computer vision community studies specific person re-identification (ReID) tasks in different scenarios separately, which limits the applications in the real world. This paper strives to resolve this problem by proposing a novel instruct-ReID t
Jun Zhang, Jiacheng Lu, Jingjing Zhang, Yu Han
Extra large-scale multiple-input multiple-output (XL-MIMO) is a key technology for future wireless communication systems. This paper considers the effects of visibility region (VR) at the base station (BS) in a non-stationary multi-user XL-MIMO scenario, where only partial antennas can receive users' signal. In time division duplexing (TDD) mode, we first es
Enhancing Road Safety: Real-Time Detection of Driver Distraction through Convolutional Neural Networks
cs.CVAmaan Aijaz Sheikh, Imaad Zaffar Khan
As we navigate our daily commutes, the threat posed by a distracted driver is at a large, resulting in a troubling rise in traffic accidents. Addressing this safety concern, our project harnesses the analytical power of Convolutional Neural Networks (CNNs), with a particular emphasis on the well-established models VGG16 and VGG19. These models are acclaimed
Kensuke Sakamoto
This paper addresses the sample selection problem in panel dyadic regression analysis. Dyadic data often include many zeros in the main outcomes due to the underlying network formation process. This not only contaminates popular estimators used in practice but also complicates the inference due to the dyadic dependence structure. We extend Kyriazidou (1997)'
Effect of insulator end cap thickness on time-dependent Hartmann flow in a rotating mirror
physics.plasm-phRahul Gaur, Ian G. Abel, Bindesh Tripathi, Egemen Kolemen
We present a framework for analyzing plasma flow in a rotating mirror. By making a series of physical assumptions, we reduce the magnetohydrodynamic (MHD) equations in a three-dimensional cylindrical system to a one-dimensional system in a shallow, cuboidal channel within a transverse magnetic field, similar to the Hartmann flow in the ducts. We then solve t