November 2024 arXiv papers — page 51
Showing 5,001–5,100 of 19,800 papers
Joshua Akin, Yunlei Zhao, Paul G. Kwiat, Elizabeth A. Goldschmidt
Quantum networking protocols, including quantum teleportation and entanglement swapping, use linear-optical Bell state measurements for heralding the distribution and transfer of quantum information. However, a linear-optical Bell state measurement requires identical photons and is susceptible to errors caused by multiphoton emission, fundamentally limiting
ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal Guidance
cs.CVHaijie Yang, Zhenyu Zhang, Hao Tang, Jianjun Qian
Diffusion models have shown impressive potential on talking head generation. While plausible appearance and talking effect are achieved, these methods still suffer from temporal, 3D or expression inconsistency due to the error accumulation and inherent limitation of single-image generation ability. In this paper, we propose ConsistentAvatar, a novel framewor
Zuyao Chen, Jinlin Wu, Zhen Lei, Chang Wen Chen
While text-to-image generation has been extensively studied, generating images from scene graphs remains relatively underexplored, primarily due to challenges in accurately modeling spatial relationships and object interactions. To fill this gap, we introduce Scene-Bench, a comprehensive benchmark designed to evaluate and enhance the factual consistency in g
Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
cs.CVZhangqi Jiang, Junkai Chen, Beier Zhu, Tingjin Luo
Hallucinations in Large Vision-Language Models (LVLMs) significantly undermine their reliability, motivating researchers to explore the causes of hallucination. However, most studies primarily focus on the language aspect rather than the visual. In this paper, we address how LVLMs process visual information and whether this process causes hallucination. Firs
Katherine Goldman
The 2-dimensional Shephard groups are quotients of 2-dimensional Artin groups by powers of standard generators. We show that such a quotient is not $\mathrm{CAT}(0)$ if the powers taken are sufficiently large. However, for a given 2-dimensional Shephard group, we construct a $\mathrm{CAT}(0)$ piecewise Euclidean cell complex with a cocompact action (analogou
Xiangtong Wang
This study introduces a new framework for analyzing capacity dynamics and throughput performance in Low Earth Orbit satellite networks (LSNs). It focuses on addressing critical gaps in existing models, particularly those concerning unreliable ISLs. Our work systematically resolves two inherent deficiencies in prior research: (1) the conflation of network cap
Qizhou Chen, Chengyu Wang, Dakan Wang, Taolin Zhang
Model editing aims to correct inaccurate knowledge, update outdated information, and incorporate new data into Large Language Models (LLMs) without the need for retraining. This task poses challenges in lifelong scenarios where edits must be continuously applied for real-world applications. While some editors demonstrate strong robustness for lifelong editin
Minoru Hirose, Hideki Murahara, Shingo Saito
The Ohno relation is one of the most celebrated results in the theory of multiple zeta values, which are iterated integrals from $0$ to $1$. In a previous paper, the authors generalized the Ohno relation to regularized multiple zeta values, which are non-admissible iterated integrals from $0$ to $1$. Meanwhile, Takeyama proved an analogue of the Ohno relatio
The Landscape of Data Reuse in Interactive Information Retrieval: Motivations, Sources, and Evaluation of Reusability
cs.IRTianji Jiang, Wenqi Li, Jiqun Liu
Sharing and reusing research data can effectively reduce redundant efforts in data collection and curation, especially for small labs and research teams conducting human-centered system research, and enhance the replicability of evaluation experiments. Building a sustainable data reuse process and culture relies on frameworks that encompass policies, standar
Shaoyang Zhou
We study the dynamics of area-preserving maps in a non-compact setting. We show that the $C^{\infty}$-closing lemma holds for area-preserving diffeomorphisms on a closed surface with finitely many points removed. As a corollary, a $C^{\infty}$-generic area-preserving diffeomorphism on such a surface has a dense set of periodic points. For area-preserving map
Yunlei Liang, Jiawei Zhu, Wen Ye, Song Gao
Spatial networks are useful for modeling geographic phenomena where spatial interaction plays an important role. To analyze the spatial networks and their internal structures, graph-based methods such as community detection have been widely used. Community detection aims to extract strongly connected components from the network and reveal the hidden relation
Integrating optimal ridesharing matching into multimodal traffic model: Implications for policy and sustainable transport system
physics.soc-phYueqi Liu, Ke Han, Zhuoqian Yang, Yanghong Yu
Integrating ridesharing matching explicitly into multimodal traffic models is crucial for accurately assessing the impacts of multimodal transport (MT) on urban economic and environmental aspects. This paper integrates an optimal ridesharing matching method into a path-based deterministic day-to-day traffic assignment framework, considers match cancellations
Jiong Wu, Kuang Gong
Deformable image registration plays an essential role in various medical image tasks. Existing deep learning-based deformable registration frameworks primarily utilize convolutional neural networks (CNNs) or Transformers to learn features to predict the deformations. However, the lack of semantic information in the learned features limits the registration pe
Ming-Fong Sie, Yen-Jui Chang, Chien-Lung Lin, Ching-Ray Chang
Over 900 million Bitcoin transactions have been recorded, posing considerable challenges for machine learning in terms of computation time and maintaining prediction accuracy. We propose an innovative approach using quantum-inspired algorithms implemented with Simulated Annealing and Quantum Annealing to address the challenge of local minima in solution spac
Discrepancy in Oil Displacement Mechanisms at the Equivalent Interfacial Tensions: Differentiating Contributions from Surfactant and Nanoparticles on Interfacial Activities
physics.flu-dynSuparit Tangparitkul, Thakheru Akamine, David Harbottle, Falan Srisuriyachai
This study examines discrepancies in oil displacement mechanisms at equivalent interfacial tensions, focusing on the distinct contributions of surfactants and nanoparticles. It was hypothesized that similar interfacial activities would result in consistent displacement outcomes, while differences would reflect unique interfacial behaviors. Micromodel experim
A. V. Sarantsev, E. Klempt, K. V. Nikonov, P. Achenbach
Photoproduction of charged pions pairs off protons is studied within the invariant masses of the final state hadrons from 1.6 to 2.4 GeV at the Thomas Jefferson National Accelerator Facility with the CLAS detector. The data are included in the Bonn-Gatchina coupled-channel analysis and provide the information necessary to determine the branching fractions fo
Learning a local trading strategy: deep reinforcement learning for grid-scale renewable energy integration
cs.LGCaleb Ju, Constance Crozier
Variable renewable generation increases the challenge of balancing power supply and demand. Grid-scale batteries co-located with generation can help mitigate this misalignment. This paper explores the use of reinforcement learning (RL) for operating grid-scale batteries co-located with solar power. Our results show RL achieves an average of 61% (and up to 96
OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining
cs.CVMing Hu, Kun Yuan, Yaling Shen, Feilong Tang
Surgical practice involves complex visual interpretation, procedural skills, and advanced medical knowledge, making surgical vision-language pretraining (VLP) particularly challenging due to this complexity and the limited availability of annotated data. To address the gap, we propose OphCLIP, a hierarchical retrieval-augmented vision-language pretraining fr
Mitchell Rosser, Marc. G Carmichael
With the recent development of natural language generation models - termed as large language models (LLMs) - a potential use case has opened up to improve the way that humans interact with robot assistants. These LLMs should be able to leverage their large breadth of understanding to interpret natural language commands into effective, task appropriate and sa
Semi-supervised Single-view 3D Reconstruction via Multi Shape Prior Fusion Strategy and Self-Attention
cs.CVWei Zhoua, Xinzhe Shia, Yunfeng Shea, Kunlong Liua
In the domain of single-view 3D reconstruction, traditional techniques have frequently relied on expensive and time-intensive 3D annotation data. Facing the challenge of annotation acquisition, semi-supervised learning strategies offer an innovative approach to reduce the dependence on labeled data. Despite these developments, the utilization of this learnin
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation
cs.DCFahao Chen, Peng Li, Zicong Hong, Zhou Su
Mixture-of-Experts (MoE) is an emerging technique for scaling large models with sparse activation. MoE models are typically trained in a distributed manner with an expert parallelism scheme, where experts in each MoE layer are distributed across multiple GPUs. However, the default expert parallelism suffers from the heavy network burden due to the all-to-all
Andrew T. McNutt, Abhinav K. Adduri, Caleb N. Ellington, Monica T. Dayao
Virtual screening of small molecules against protein targets can accelerate drug discovery and development by predicting drug-target interactions (DTIs). However, structure-based methods like molecular docking are too slow to allow for broad proteome-scale screens, limiting their application in screening for off-target effects or new molecular mechanisms. Re
Hoyoung Kim, Seokhee Jin, Changhwan Sung, Jaechang Kim
Vision-language models (VLMs) have demonstrated remarkable zero-shot performance across various classification tasks. Nonetheless, their reliance on hand-crafted text prompts for each task hinders efficient adaptation to new tasks. While prompt learning offers a promising solution, most studies focus on maximizing the utilization of given few-shot labeled da
Biswajit Sahoo, Akilan K, Katherine Matthews, Alexandre Pofelski
Complex ferromagnetic oxides such as La$_{0.67}$Sr$_{0.33}$MnO$_3$ (LSMO) offer pathways for creating energy efficient spintronic devices with new functionalities. LSMO exhibits high-temperature ferromagnetism, half metallicity, sharp resonance linewidth, low damping and a large anisotropic magnetoresistance response. Combined with Pt, a proven material with
Gayatri Priyadarsini Kancherla, Dishank Goel, Abhishek Bichhawat
Web applications often include third-party content and scripts to personalize a user's online experience. These scripts have unrestricted access to a user's private data stored in the browser's persistent storage like cookies, localstorage and IndexedDB, associated with the host page. Various mechanisms have been implemented to restrict access to these stora
Anna Frebel
Ancient, long-lived stars remain present in all components of our home galaxy, the Milky Way. Born a few hundred million after the Big Bang and during a time that marked the very beginning of the chemical evolution, these stars display very low abundances of elements heavier and hydrogen and helium, making them "metal-poor". Studying the chemical composition
Gayatri Priyadarsini Kancherla, Nataliia Bielova, Cristiana Santos, Abhishek Bichhawat
The GDPR requires websites to facilitate the right to revoke consent from Web users. While numerous studies measured compliance of consent with the various consent requirements, no prior work has studied consent revocation on the Web. Therefore, it remains unclear how difficult it is to revoke consent on the websites' interfaces, nor whether revoked consent
FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
cs.CVTrong Thang Pham, Ngoc-Vuong Ho, Nhat-Tan Bui, Thinh Phan
Developing an interpretable system for generating reports in chest X-ray (CXR) analysis is becoming increasingly crucial in Computer-aided Diagnosis (CAD) systems, enabling radiologists to comprehend the decisions made by these systems. Despite the growth of diverse datasets and methods focusing on report generation, there remains a notable gap in how closel
Richard Stone
This paper starts by introducing results from geometric measure theory to prove symmetric decreasing rearrangement inequalities on $\mathbb{R}^n$, which give multiple proofs of the isoperimetric and P\'{o}lya-Szeg\H{o} inequalities. Then we consider smooth oriented Riemannian manifolds of the form $M^n = (0,\infty)\times \Sigma^{n-1}$, and test what results
Hang Hua, Qing Liu, Lingzhi Zhang, Jing Shi
The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, including image and video captioning, visual question answering, and cross-modal retrieval. Despite their superior capabilities, VLMs struggle with fine-grained image regional compositi
Minimizing Nature's Cost: Exploring Data-Free Physics-Informed Neural Network Solvers for Fluid Mechanics Applications
physics.flu-dynAbdelrahman Elmaradny, Ahmed Atallah, Haithem Taha
In this paper, we present a novel approach for fluid dynamic simulations by harnessing the capabilities of Physics-Informed Neural Networks (PINNs) guided by the newly unveiled principle of minimum pressure gradient (PMPG). In a PINN formulation, the physics problem is converted into a minimization problem (typically least squares). The PMPG asserts that for
Ilkin Aliyev, Jesus Lopez, Tosiron Adegbija
Spiking Neural Networks (SNNs) offer potential advantages in energy efficiency but currently trail Artificial Neural Networks (ANNs) in versatility, largely due to challenges in efficient input encoding. Recent work shows that direct coding achieves superior accuracy with fewer timesteps than traditional rate coding. However, there is a lack of specialized h
Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks
cs.CVHan Wang, Gang Wang, Huan Zhang
Vision Language Models (VLMs) can produce unintended and harmful content when exposed to adversarial attacks, particularly because their vision capabilities create new vulnerabilities. Existing defenses, such as input preprocessing, adversarial training, and response evaluation-based methods, are often impractical for real-world deployment due to their high
Exploring Large Language Models for Multimodal Sentiment Analysis: Challenges, Benchmarks, and Future Directions
cs.CLShezheng Song
Multimodal Aspect-Based Sentiment Analysis (MABSA) aims to extract aspect terms and their corresponding sentiment polarities from multimodal information, including text and images. While traditional supervised learning methods have shown effectiveness in this task, the adaptability of large language models (LLMs) to MABSA remains uncertain. Recent advances i
Haoyu Wu, Jingyi Xu, Hieu Le, Dimitris Samaras
Token merging can effectively accelerate various vision systems by processing groups of similar tokens only once and sharing the results across them. However, existing token grouping methods are often ad hoc and random, disregarding the actual content of the samples. We show that preserving high-information tokens during merging - those essential for semanti
Hua Qiu, Qi Wang, Shufang Wang
We calculate the Assouad and lower dimensions of graph-directed Bedford-McMullen carpets, which reflect the extreme local scaling laws of the sets, in contrasting with known results on Hausdorff and box dimensions. We also investigate the relationship between distinct dimensions. In particular, we identify an equivalent condition when the box and Assouad dim
Pengzhi Xie
For any weakly interacting particle system with bounded kernel, we give uniform-in-time estimates of the $L^2$ norm of correlation functions, provided that the diffusion coefficient is large enough. When the condition on the kernels is more restrictive, we can remove the dependence of the lower bound for diffusion coefficient on the initial data and estimate
ML-SPEAK: A Theory-Guided Machine Learning Method for Studying and Predicting Conversational Turn-taking Patterns
cs.CLLisa R. O'Bryan, Madeline Navarro, Juan Segundo Hevia, Santiago Segarra
Predicting team dynamics from personality traits remains a fundamental challenge for the psychological sciences and team-based organizations. Understanding how team composition generates team processes can significantly advance team-based research along with providing practical guidelines for team staffing and training. Although the Input-Process-Output (IPO
A Comparative Analysis of Transformer and LSTM Models for Detecting Suicidal Ideation on Reddit
cs.LGKhalid Hasan, Jamil Saquer
Suicide is a critical global health problem involving more than 700,000 deaths yearly, particularly among young adults. Many people express their suicidal thoughts on social media platforms such as Reddit. This paper evaluates the effectiveness of the deep learning transformer-based models BERT, RoBERTa, DistilBERT, ALBERT, and ELECTRA and various Long Short
Partial Knowledge Distillation for Alleviating the Inherent Inter-Class Discrepancy in Federated Learning
cs.LGXiaoyu Gan, Jingbo Jiang, Jingyang Zhu, Xiaomeng Wang
Substantial efforts have been devoted to alleviating the impact of the long-tailed class distribution in federated learning. In this work, we observe an interesting phenomenon that certain weak classes consistently exist even for class-balanced learning. These weak classes, different from the minority classes in the previous works, are inherent to data and r
Scale-invariant total decay width $\Gamma(H\to b\bar{b})$ using the novel method of characteristic operator
hep-phJiang Yan, Xing-Gang Wu, Jian-Ming Shen, Xu-Dong Huang
In this paper, a novel method via using the characteristic operator~(CO) ${\cal \hat{D}}_{n_{\gamma}, n_{\beta}}$ is proposed to extend the applicability of PMC, which is a theoretical generalization of previous PMC single-scale setting approach. Using the CO formulism, we are able to facilitate the derivation of complex scenarios within a structured theoret
Ruodu Wang, Qinyu Wu
Given two random variables taking values in a bounded interval, we study whether one dominates the other in higher-order stochastic dominance depends on the reference interval in the model setting. We obtain two results. First, the stochastic dominance relations get strictly stronger when the reference interval shrinks if and only if the order of stochastic
Dylan Clark-Boucher, Brent A Coull, Harrison T Reeder, Fenglei Wang
A key challenge in differential abundance analysis of microbial samples is that the counts for each sample are compositional, resulting in biased comparisons of the absolute abundance across study groups. Normalization-based differential abundance analysis methods rely on external normalization factors that account for the compositionality by standardizing t
Xiaoling Hu, Xiangrui Zeng, Oula Puonti, Juan Eugenio Iglesias
Domain randomization through synthesis is a powerful strategy to train networks that are unbiased with respect to the domain of the input images. Randomization allows networks to see a virtually infinite range of intensities and artifacts during training, thereby minimizing overfitting to appearance and maximizing generalization to unseen data. Although powe
Varatheepan Paramanayakam, Andreas Karatzas, Iraklis Anagnostopoulos, Dimitrios Stamoulis
The advanced function-calling capabilities of foundation models open up new possibilities for deploying agents to perform complex API tasks. However, managing large amounts of data and interacting with numerous APIs makes function calling hardware-intensive and costly, especially on edge devices. Current Large Language Models (LLMs) struggle with function ca
Stanley E. Lazic
Statistical inference often conflates the probability of a parameter with the probability of a hypothesis, a critical misunderstanding termed the ultimate issue error. This error is pervasive across the social, biological, and medical sciences, where null hypothesis significance testing (NHST) is mistakenly understood to be testing hypotheses rather than eva
Leonidas Gee, Wing Yan Li, Viktoriia Sharmanska, Novi Quadrianto
The cost of deploying vision transformers increasingly represents a barrier to wider industrial adoption. Existing compression techniques require additional end-to-end fine-tuning or incur a significant drawback to energy efficiency, making them ill-suited for online (real-time) inference, where a prediction is made on any new input as it comes in. We introd
The Decoy Dilemma in Online Medical Information Evaluation: A Comparative Study of Credibility Assessments by LLM and Human Judges
cs.IRJiqun Liu, Jiangen He
Can AI be cognitively biased in automated information judgment tasks? Despite recent progresses in measuring and mitigating social and algorithmic biases in AI and large language models (LLMs), it is not clear to what extent LLMs behave "rationally", or if they are also vulnerable to human cognitive bias triggers. To address this open problem, our study, con
ChatBCI: A P300 Speller BCI Leveraging Large Language Models for Improved Sentence Composition in Realistic Scenarios
cs.HCJiazhen Hong, Weinan Wang, Laleh Najafizadeh
P300 speller BCIs allow users to compose sentences by selecting target keys on a GUI through the detection of P300 component in their EEG signals following visual stimuli. Most P300 speller BCIs require users to spell words letter by letter, or the first few initial letters, resulting in high keystroke demands that increase time, cognitive load, and fatigue.
Manami Roy, Smita Mathur, Sanskriti Das, Armando Lara-DI
Recent observations have revealed a super-virial temperature gas phase at log(T/K) $\sim7$ in the Milky Way, challenging existing galaxy-formation models. This hot gas phase was discovered toward extragalactic absorption sightlines and blank-sky emission fields, both at high galactic latitudes. The location of this hot component is unknown; is it in the exte
Faisal N. Abu-Khzam, Lucas Isenmann, Sergio Thoumi
Correlation clustering seeks a partition of the vertex set of a given graph/network into groups of closely related, or just close enough, vertices so that elements of different groups are not close to each other. The problem has been previously modeled and studied as a graph editing problem, namely Cluster Editing, which assumes that closely related data ele
From Quantum Cognition to Conceptuality Interpretation I: Tracing the Brussels Group's Intellectual Journey
physics.hist-phDiederik Aerts, Massimiliano Sassoli de Bianchi, Sandro Sozzo
The conceptuality interpretation of quantum mechanics proposes that quantum entities have a conceptual nature, interacting with the material world through processes that are the physical counterpart of the meaning-based processes which typically occur in human cognition. This interpretation emerged from the early developments in quantum cognition, a field th
Rahul Shenoy, Zhihong Pan, Kaushik Balakrishnan, Qisen Cheng
Image generation using diffusion models have demonstrated outstanding learning capabilities, effectively capturing the full distribution of the training dataset. They are known to generate wide variations in sampled images, albeit with a trade-off in image fidelity. Guided sampling methods, such as classifier guidance (CG) and classifier-free guidance (CFG),
Vladimir Sluchak
There is a class of physical filtration processes where the input is adequately modeled by a continuous periodic function f (x) of bounded variation over its period, and the output depends only on certain harmonics of the Fourier expansion of f (x) in the orthogonal basis of trigonometric functions. One example is the discrete spectrum sound generation by a
Decoding physics identity: A Spanish-language adaptation on an instrument and its correlation with STEM achievement
physics.ed-phO. I. González-Peña, G. Morán-Soto, B. M. Rodríguez-Lara
The representation of Spanish-speaking students in STEM identity literature, particularly in physics identity, is conspicuously minimal. This study addresses this gap with a two-pronged approach. First, a physics identity instrument was adapted for Spanish-speaking STEM students to promote inclusivity in educational research. Data from 334 Mexican STEM stude
James Conley, Konstantinos Georgiou
We consider $n$ unit-speed mobile agents initially positioned at the center of a unit disk, tasked with inspecting all points on the disk's perimeter. A perimeter point is considered covered if an agent located outside the disk's interior has unobstructed visibility of it, treating the disk itself as an obstacle. For $n=1$, this problem is known as the shore
The Hatching-Box: A Novel System for Automated Monitoring and Quantification of Drosophila melanogaster Developmental Behavior
cs.CVJulian Bigge, Maite Ogueta, Luis Garcia, Benjamin Risse
In this paper we propose the Hatching-Box, a novel imaging and analysis system to automatically monitor and quantify the developmental behavior of Drosophila in standard rearing vials and during regular rearing routines, rendering explicit experiments obsolete. This is achieved by combining custom tailored imaging hardware with dedicated detection and tracki
Xiaojun Chen, Jieheng Zeng
Let $V$ be a finite dimensional $k$-vector space, where $k$ is an algebraic closed field of characteristic zero. Let $G \subseteq \mathrm{SL}(V)$ be a finite abelian group, and denote by $S$ the $G$-invariant subring of the polynomial ring $k[V]$. It is shown that the singularity category $D_{sg}(S)$ recovers the reduced singular locus of $\mathrm{Spec}(S)$.
Chiara Mauri, Ryan Fritz, Jocelyn Mora, Benjamin Billot
The claustrum is a band-like gray matter structure located between putamen and insula whose exact functions are still actively researched. Its sheet-like structure makes it barely visible in in vivo Magnetic Resonance Imaging (MRI) scans at typical resolutions and neuroimaging tools for its study, including methods for automatic segmentation, are currently v
Mara Finkelstein, Dan Deutsch, Parker Riley, Juraj Juraska
As LLMs continue to become more powerful and versatile, human evaluation has quickly become intractable at scale and reliance on automatic metrics has become the norm. Recently, it has been shown that LLMs are themselves state-of-the-art evaluators for many tasks. These Autoraters are typically designed so that they generalize to new systems and test sets. I
Artem Karpov, Seong Hah Cho, Austin Meek, Raymond Koopmanschap
In this work, we study the alignment (BrainScore) of large language models (LLMs) fine-tuned for moral reasoning on behavioral data and/or brain data of humans performing the same task. We also explore if fine-tuning several LLMs on the fMRI data of humans performing moral reasoning can improve the BrainScore. We fine-tune several LLMs (BERT, RoBERTa, DeBERT
Arif Kerem Dayi, Sitan Chen
LoRA has emerged as one of the de facto methods for fine-tuning foundation models with low computational cost and memory footprint. The idea is to only train a low-rank perturbation to the weights of a pre-trained model, given supervised data for a downstream task. Despite its empirical sucess, from a mathematical perspective it remains poorly understood wha
Antonio Casares, Olivier Idir, Denis Kuperberg, Corto Mascle
We present a polynomial-time algorithm minimising the number of states of history-deterministic generalised coBüchi automata, building on the work of Abu Radi and Kupferman on coBüchi automata. On the other hand, we establish that the minimisation problem for both deterministic and history-deterministic generalised Büchi automata is NP-complete, as well as t
Julien Bect, Niklas Georg, Ulrich Römer, Sebastian Schöps
This work is concerned with the kernel-based approximation of a complex-valued function from data, where the frequency response function of a partial differential equation in the frequency domain is of particular interest. In this setting, kernel methods are employed more and more frequently, however, standard kernels do not perform well. Moreover, the role
Vedran Vujnović, Nenad Kralj, Marin Karuza
In this paper, we theoretically analyze the optimization of a Fabry-P\'{e}rot cavity for the purpose of detecting partially absorbing objects placed inside without photon exchange. Utilizing the input-output formalism, we quantitatively relate the probability of correctly inferring the presence or absence of the object to the probability of avoiding absorpti
Magnetic ground state of the dimer-based hexagonal perovskite Ba$_{3}$ZnRu$_{2}$O$_{9}$
cond-mat.str-elS. Hayashida, H. Gretarsson, P. Puphal, M. Isobe
We investigate the magnetic ground state of single crystals of the ruthenium-dimer-based hexagonal perovskite Ba$_{3}$ZnRu$_{2}$O$_{9}$ using magnetic susceptibility and resonant inelastic x-ray scattering (RIXS) measurements. While a previous study on powder samples exhibited intriguing magnetic behavior, questions about whether the spin state within a Ru$_
S P Sharan, Minkyu Choi, Sahil Shah, Harsh Goel
Recent advancements in text-to-video models such as Sora, Gen-3, MovieGen, and CogVideoX are pushing the boundaries of synthetic video generation, with adoption seen in fields like robotics, autonomous driving, and entertainment. As these models become prevalent, various metrics and benchmarks have emerged to evaluate the quality of the generated videos. How
Elita Lobo, Chirag Agarwal, Himabindu Lakkaraju
Large language models have emerged as powerful tools for general intelligence, showcasing advanced natural language processing capabilities that find applications across diverse domains. Despite their impressive performance, recent studies have highlighted the potential for significant enhancements in LLMs' task-specific performance through fine-tuning strat
Sohaib Ahmad, Qizheng Yang, Haoliang Wang, Ramesh K. Sitaraman
Text-to-image generation using diffusion models has gained increasing popularity due to their ability to produce high-quality, realistic images based on text prompts. However, efficiently serving these models is challenging due to their computation-intensive nature and the variation in query demands. In this paper, we aim to address both problems simultaneou
Hao Liu
Deep learning models often require specially designed architectures to process data of different dimensions, such as 1D time series, 2D images, and 3D volumetric data. Existing bidirectional models mainly focus on sequential data, making it difficult to scale effectively to higher dimensions. To address this issue, we propose a novel multi-dimensional bidire
Zhirayr Avetisyan, Alexey Karapetyants
We describe a general framework of functional and Fourier analysis on domains with a free action of an Abelian Lie group $G$. Namely, on a domain of the form $G\times Y$ we introduce the appropriate spaces of distributions and measurable functions, establishing their most basic properties. Then we consider the half-Fourier transform $f(x,y)\mapsto\hat f(\xi,
Scout Jarman, Zigfried Hampel-Arias, Adra Carr, Kevin R. Moon
Longwave infrared (LWIR) hyperspectral imaging can be used for many tasks in remote sensing, including detecting and identifying effluent gases by LWIR sensors on airborne platforms. Once a potential plume has been detected, it needs to be identified to determine exactly what gas or gases are present in the plume. During identification, the background undern
Simone Franchini
We discuss the relation between p-adic numbers and kernels in view of a recent large deviation theory for mean-field spin glasses. As an application we show several fundamental properties of numerical bases in kernel language. In particular, we show that the Derrida's Generalized Random Energy Model can be interpreted as a (random) numerical base. We also sh
Zhen-Ni Xu, Daniele Binosi, Chen Chen, Khépani Raya
Using available information from Drell-Yan data on pion and kaon structure functions, an approach is described which enables the development of pointwise profiles for all pion and kaon parton distribution functions (DFs) without reference to theories of hadron structure. The key steps are construction of structure-function-constrained probability-weighted en
Ilia Zaznov, Atta Badii, Alfonso Dufour, Julian Kunkel
AdamZ is an advanced variant of the Adam optimiser, developed to enhance convergence efficiency in neural network training. This optimiser dynamically adjusts the learning rate by incorporating mechanisms to address overshooting and stagnation, that are common challenges in optimisation. Specifically, AdamZ reduces the learning rate when overshooting is dete
Kevin J Napier
Numerical solutions of Kepler's Equation are critical components of celestial mechanics software, and are often computation hot spots. This work uses symbolic regression and a genetic learning algorithm to find new initial guesses for iterative Kepler solvers for both elliptical and hyperbolic orbits. The new initial guesses are simple to implement, and resu
Alexandre P. Costa, Lucas Queiroz, Danilo T. Alves
Recently, it has been shown that, under the action of the lateral van der Waals (vdW) force due to a perfectly conducting corrugated plane, a neutral anisotropic polarizable particle in vacuum can be attracted not only to the nearest corrugation peak but also to a valley or an intermediate point between a peak and a valley, with such behaviors called the pea
Whispering-Gallery-Mode Resonators for Detection and Classification of Free-Flowing Nanoparticles and Cells through Photoacoustic Signatures
physics.bio-phJie Liao, Maxwell Adolphson, Hangyue Li, Dipayon Kumar Sikder
Micro and nanoscale particles are crucial in various fields, from biomedical imaging to environmental processes. While conventional spectroscopy and microscopy methods for characterizing these particles often involve bulky equipment and complex sample preparation, optical micro-sensors have emerged as a promising alternative. However, their broad applicabili
Transforming NLU with Babylon: A Case Study in Development of Real-time, Edge-Efficient, Multi-Intent Translation System for Automated Drive-Thru Ordering
cs.CLMostafa Varzaneh, Pooja Voladoddi, Tanmay Bakshi, Uma Gunturi
Real-time conversational AI agents face challenges in performing Natural Language Understanding (NLU) in dynamic, outdoor environments like automated drive-thru systems. These settings require NLU models to handle background noise, diverse accents, and multi-intent queries while operating under strict latency and memory constraints on edge devices. Additiona
Mani Amani, Reza Akhavian
Construction robots have gained significant traction in recent years in research and development. However, the application of industrial robots has unique challenges. Dynamic environments, domain-specific tasks, and complex localization and mapping are significant obstacles in their development. In construction job sites, moving objects and complex machinery
Gautham Vasan, Mohamed Elsayed, Alireza Azimi, Jiamin He
Modern deep policy gradient methods achieve effective performance on simulated robotic tasks, but they all require large replay buffers or expensive batch updates, or both, making them incompatible for real systems with resource-limited computers. We show that these methods fail catastrophically when limited to small replay buffers or during incremental lear
Peter Eisenberger, Matthew Realff
It is now accepted that gigatonnes of Carbon Dioxide Removal (CDR) from the atmosphere are needed to avoid the threat of catastrophic climate change. Direct Air Capture (DAC) is a promising scalable CDR with a relatively small environmental footprint. But questions about DAC cost and energy use remain that are delaying the needed DAC policy decisions to crea
The Power of Types: Exploring the Impact of Type Checking on Neural Bug Detection in Dynamically Typed Languages
cs.SEBoqi Chen, José Antonio Hernández López, Gunter Mussbacher, Dániel Varró
Motivation: Automated bug detection in dynamically typed languages such as Python is essential for maintaining code quality. The lack of mandatory type annotations in such languages can lead to errors that are challenging to identify early with traditional static analysis tools. Recent progress in deep neural networks has led to increased use of neural bug d
Exploiting Watermark-Based Defense Mechanisms in Text-to-Image Diffusion Models for Unauthorized Data Usage
cs.CVSoumil Datta, Shih-Chieh Dai, Leo Yu, Guanhong Tao
Text-to-image diffusion models, such as Stable Diffusion, have shown exceptional potential in generating high-quality images. However, recent studies highlight concerns over the use of unauthorized data in training these models, which may lead to intellectual property infringement or privacy violations. A promising approach to mitigate these issues is to app
Personalization of Wearable Sensor-Based Joint Kinematic Estimation Using Computer Vision for Hip Exoskeleton Applications
cs.ROChangseob Song, Bogdan Ivanyuk-Skulskyi, Adrian Krieger, Kaitao Luo
Accurate lower-limb joint kinematic estimation is critical for applications such as patient monitoring, rehabilitation, and exoskeleton control. While previous studies have employed wearable sensor-based deep learning (DL) models for estimating joint kinematics, these methods often require extensive new datasets to adapt to unseen gait patterns. Meanwhile, r
Hongzhi Liu, Zhizhang Xie, Guoliang Yu
For each orientation-preserving homotopy equivalence between two closed oriented smooth manifolds, there are mainly two different approaches to the higher $\rho$ invariant associated to this homotopy equivalence. In this article, we show that these two definitions of the higher $\rho$ invariant are equivalent.
Sebastian Siebertz, Alexandre Vigny
Tractability results for the model checking problem of logics yield powerful algorithmic meta theorems of the form: Every computational problem expressible in a logic $L$ can be solved efficiently on every class $\mathscr{C}$ of structures satisfying certain conditions. The most prominent logics studied in the field are (counting) monadic second-order logic
Moses Charikar, Chirag Pabbaraju
The recent work of Kleinberg & Mullainathan [KM24] provides a concrete model for language generation in the limit: given a sequence of examples from an unknown target language, the goal is to generate new examples from the target language such that no incorrect examples are generated beyond some point. In sharp contrast to strong negative results for the clo
Unwanted couplings can induce amplification in quantum memories despite negligible apparent noise
quant-phFaezeh Kimiaee Asadi, Janish Kumar, Jiawei Ji, Khabat Heshami
Theoretical quantum memory design often involves selectively focusing on certain energy levels to mimic an ideal $\Lambda$-configuration, a common approach that may unintentionally overlook the impact of neighboring levels or undesired couplings. While this simplification may be justified in certain protocols or platforms, it can significantly distort the ac
Boosting Photon-Number-Resolved Detection Rates of Transition-Edge Sensors by Machine Learning
quant-phZhenghao Li, Matthew J. H. Kendall, Gerard J. Machado, Ruidi Zhu
Transition-Edge Sensors (TESs) are very effective photon-number-resolving (PNR) detectors that have enabled many photonic quantum technologies. However, their relatively slow thermal recovery time severely limits their operation rate in experimental scenarios compared to leading non-PNR detectors. In this work, we develop an algorithmic approach that enables
PaRCE: Probabilistic and Reconstruction-based Competency Estimation for CNN-based Image Classification
cs.CVSara Pohland, Claire Tomlin
Convolutional neural networks (CNNs) are extremely popular and effective for image classification tasks but tend to be overly confident in their predictions. Various works have sought to quantify uncertainty associated with these models, detect out-of-distribution (OOD) inputs, or identify anomalous regions in an image, but limited work has sought to develop
Nivetha Jayakumar, Srivardhan Reddy Gadila, Tonmoy Hossain, Yangfeng Ji
Preserving topological structures is important in real-world applications, particularly in sensitive domains such as healthcare and medicine, where the correctness of human anatomy is critical. However, most existing image editing models focus on manipulating intensity and texture features, often overlooking object geometry within images. To address this iss
Francesco Genovese, Wendy Lowen, Julie Symons, Michel Van den Bergh
This paper provides the final ingredient in the development of the deformation theory of pretriangulated dg-categories endowed with a nice t-structure, which was initiated by the authors and is modeled after the previously developed deformation theory of abelian categories. We show how to extend a t-structure on a pretriangulated dg-category to its dg-derive
Imed Basdouri, Bouzid Mosbahi
We describe Rota-Baxter operators, Reynolds operators, Nijenhuis operators, and Averaging operators on 2-dimensional dendriform algebras over $\mathbb{C}$.
Chamani M. Gunasekera, Peter A. M. van Hoof, Masahiro Tsujimoto, Gary J. Ferland
We present a simple, yet powerful column density diagnostic for plasmas enabled by X-ray microcalorimeter observations. With the recent developments of the spectral simulation code Cloudy, inspired by the high spectral resolution of the X-Ray Imaging and Spectroscopy Mission (XRISM) and the Advanced Telescope for High Energy Astrophysics (Athena), we make pr
Regulator-Manufacturer AI Agents Modeling: Mathematical Feedback-Driven Multi-Agent LLM Framework
cs.AIYu Han, Zekun Guo
The increasing complexity of regulatory updates from global authorities presents significant challenges for medical device manufacturers, necessitating agile strategies to sustain compliance and maintain market access. Concurrently, regulatory bodies must effectively monitor manufacturers' responses and develop strategic surveillance plans. This study em
UniGaussian: Driving Scene Reconstruction from Multiple Camera Models via Unified Gaussian Representations
cs.CVYuan Ren, Guile Wu, Runhao Li, Zheyuan Yang
Urban scene reconstruction is crucial for real-world autonomous driving simulators. Although existing methods have achieved photorealistic reconstruction, they mostly focus on pinhole cameras and neglect fisheye cameras. In fact, how to effectively simulate fisheye cameras in driving scene remains an unsolved problem. In this work, we propose UniGaussian, a
Hasan B. Al Ba'ba'a
This article introduces a methodology for inducing wavenumber bandgaps via alternating Willis coupling signs. A non-reciprocal wave equation of Willis-type is first considered, and its wave dispersion analyses are carried out via the transfer matrix method. By creating unit cells from two identical Willis-type elastic layers, yet with reversed Willis-couplin
Zhuoran Tan, Christos Anagnostopoulos, Shameem P. Parambath, Jeremy Singer
Multi-source logs provide a comprehensive overview of ongoing system activities, allowing for in-depth analysis to detect potential threats. A practical approach for threat detection involves explicit extraction of entity triples (subject, action, object) towards building provenance graphs to facilitate the analysis of system behavior. However, current log p
Taewook Kim, Ze Wang, Zhengyuan Yang, Jiang Wang
Text-to-image diffusion models have demonstrated tremendous success in synthesizing visually stunning images given textual instructions. Despite remarkable progress in creating high-fidelity visuals, text-to-image models can still struggle with precisely rendering subjects, such as text spelling. To address this challenge, this paper explores using additiona