October 2024 arXiv papers — page 86
Showing 8,501–8,600 of 23,665 papers
Terne Sasha Thorn Jakobsen, Andreas Bjerre-Nielsen, Robert Böhm
Crowdsourced annotations of data play a substantial role in the development of Artificial Intelligence (AI). It is broadly recognised that annotations of text data can contain annotator bias, where systematic disagreement in annotations can be traced back to differences in the annotators' backgrounds. Being unaware of such annotator bias can lead to represen
Large Deviation Theory Approach to Fluctuation Theorems and Landauer's Principle through Heat Redefinition
physics.chem-phTatsuaki Tsuruyama
Large deviation theory (LDT) provides a mathematical framework to quantify the probabilities of rare events in stochastic systems. In this study, we applied LDT to model a chemical reaction system and demonstrated that the fluctuation theorem for nonequilibrium reaction systems can be derived from the symmetry of the cumulant generating function defined thro
Zifan Peng, Yingjie Xue, Jingyu Liu
Options are fundamental to blockchain-based financial services, offering essential tools for risk management and price speculation, which enhance liquidity, flexibility, and market efficiency in decentralized finance (DeFi). Despite the growing interest in options for blockchain-resident assets, such as cryptocurrencies, current option mechanisms face signif
Shpresim Sadiku, Moritz Wagner, Sai Ganesh Nagarajan, Sebastian Pokutta
We study the problem of finding optimal sparse, manifold-aligned counterfactual explanations for classifiers. Canonically, this can be formulated as an optimization problem with multiple non-convex components, including classifier loss functions and manifold alignment (or \emph{plausibility}) metrics. The added complexity of enforcing \emph{sparsity}, or sho
Corentin Faipeur
In this paper, we study a model of long-range site percolation on graphs of bounded degree, namely the Boolean percolation model. In this model, each vertex of an infinite connected graph is the center of a ball of random radius, and vertices are said to be active independently with probability $p \in [0, 1]$. We consider $W$ to be the reunion of random ball
Raphaël Carpintero Perez, Sébastien da Veiga, Josselin Garnier, Brian Staber
In computational physics, machine learning has now emerged as a powerful complementary tool to explore efficiently candidate designs in engineering studies. Outputs in such supervised problems are signals defined on meshes, and a natural question is the extension of general scalar output regression models to such complex outputs. Changes between input geomet
Efficient Non-Myopic Layered Bayesian Optimization For Large-Scale Bathymetric Informative Path Planning
cs.ROAlexander Kiessling, Ignacio Torroba, Chelsea Rose Sidrane, Ivan Stenius
Informative path planning (IPP) applied to bathymetric mapping allows AUVs to focus on feature-rich areas to quickly reduce uncertainty and increase mapping efficiency. Existing methods based on Bayesian optimization (BO) over Gaussian Process (GP) maps work well on small scenarios but they are short-sighted and computationally heavy when mapping larger area
Ziwei Zhao, Xiangmei Ma, Paul Milligan, Yin Bun Cheung
Background: The Cox model and its extensions assuming proportional hazards is widely used to estimate vaccine efficacy (VE). In the typical situation that VE wanes over time, the VE estimates are not only sensitive to study duration and timing of vaccine delivery in relation to disease seasonality but also biased in the presence of sample attrition. Furtherm
Tianhang Lu
The goal of this paper is to establish a decomposition of the network based on the maximum flow problem.
Frank Nielsen
An inductive mean is a mean defined as a limit of a convergence sequence of other means. Historically, this notion of inductive means obtained as limits of sequences was pioneered independently by Lagrange and Gauss for defining the arithmetic-geometric mean. In this note, we first explain several generalizations of the scalar geometric mean to symmetric pos
Xinyu Yuan, Yan Qiao, Pei Zhao, Rongyao Hu
The traffic matrix estimation (TME) problem has been widely researched for decades of years. Recent progresses in deep generative models offer new opportunities to tackle TME problems in a more advanced way. In this paper, we leverage the powerful ability of denoising diffusion probabilistic models (DDPMs) on distribution learning, and for the first time ado
Andrii Rohovyi, Peter J. Stuckey, Toby Walsh
Faster pathfinding in time-dependent transport networks is an important and challenging problem in navigation systems. There are two main types of transport networks: road networks for car driving and public transport route network. The solutions that work well in road networks, such as Time-dependent Contraction Hierarchies and other graph-based approaches,
Imanol Echeverria, Maialen Murua, Roberto Santana
Recent advances in deep learning have shown significant potential for solving combinatorial optimization problems in real-time. Unlike traditional methods, deep learning can generate high-quality solutions efficiently, which is crucial for applications like routing and scheduling. However, existing approaches like deep reinforcement learning (RL) and behavio
R. Alsulami, S. Einecke, G. P. Rowell, P. K. McGee
We investigate the unusual H$\alpha$ features found towards the Scutum Supershell via recent arc-minute and arc-second resolution imaging. These multi-degree features resemble a long central spine ending in a bow-shock morphology. We performed a multi-wavelength study in [SII] optical, radio continuum, infrared continuum, HI, CO, X-ray and gamma-ray emission
Quantiles and Quantile Regression on Riemannian Manifolds: a measure-transportation-based approach
math.STMarc Hallin, Hang Liu
Increased attention has been given recently to the statistical analysis of variables with values on nonlinear manifolds. A natural but nontrivial problem in that context is the definition of quantile concepts. We are proposing a solution for compact Riemannian manifolds without boundaries; typical examples are polyspheres, hyperspheres, and toro\"{\i}dal man
Yuchen Wu, Yifan Yang, Gang Xu, Junjie Cao
Cooperative path planning, a crucial aspect of multi-agent systems research, serves a variety of sectors, including military, agriculture, and industry. Many existing algorithms, however, come with certain limitations, such as simplified kinematic models and inadequate support for multiple group scenarios. Focusing on the planning problem associated with a n
Xiangjian Qian, Jiale Huang, Mingpu Qin
Recent studies have highlighted the combination of tensor network methods and the stabilizer formalism as a very effective framework for simulating quantum many-body systems, encompassing areas from ground state to time evolution simulations. In these approaches, the entanglement associated with stabilizers is transferred to Clifford circuits, which can be e
Falko Schmidt, Carlos David Gonzalez-Gomez, Emilio Ruiz-Reina, Raul A. Rica
Microfluidics has revolutionized control over small volumes through the use of physical barriers. However, the rigidity of these barriers limits flexibility in applications. We present an optofluidic toolbox that leverages structured light and photothermal conversion to create dynamic, reconfigurable fluidic boundaries. This system enables precise manipulati
Changes in Sentiments and User Engagement for 2024 U.S. Presidential Candidates After Biden's Withdrawal: An Analysis of TikTok Videos
cs.SIYuwei Chuai, Gabriele Lenzini
The 2024 U.S. presidential election has sparked widespread online discussions about the presidential candidates. Joe Biden's withdrawal from the race and Kamala Harris's subsequent entry as the Democratic candidate likely alter the dynamics of these online discussions; yet, this hypothesis requires evidence. Here, we study how sentiments and user engagement
Estimating Individual Dose-Response Curves under Unobserved Confounders from Observational Data
cs.LGShutong Chen, Yang Li
Estimating an individual's potential response to continuously varied treatments is crucial for addressing causal questions across diverse domains, from healthcare to social sciences. However, existing methods are limited either to estimating causal effects of binary treatments, or scenarios where all confounding variables are measurable. In this work, we pre
Ankur Kumar
KV cache compression methods have mainly relied on scalar quantization techniques to reduce the memory requirements during decoding. In this work, we apply residual vector quantization, which has been widely used for high fidelity audio compression, to compress KV cache in large language models (LLM). We adapt the standard recipe with minimal changes to comp
Jan Ebr
The Pierre Auger Observatory, the world's largest observatory of ultra-high-energy cosmic rays (UHECR), offers a unique insight into the properties of hadronic interactions occurring in air showers at energies well above those reached at human-made accelerators. The key probe into the hadronic interactions has, for a long time, been the number of muons arriv
Mitigating Hallucinations of Large Language Models in Medical Information Extraction via Contrastive Decoding
cs.CLDerong Xu, Ziheng Zhang, Zhihong Zhu, Zhenxi Lin
The impressive capabilities of large language models (LLMs) have attracted extensive interests of applying LLMs to medical field. However, the complex nature of clinical environments presents significant hallucination challenges for LLMs, hindering their widespread adoption. In this paper, we address these hallucination issues in the context of Medical Infor
When LLMs Learn to be Students: The SOEI Framework for Modeling and Evaluating Virtual Student Agents in Educational Interaction
cs.CVYiping Ma, Shiyu Hu, Xuchen Li, Yipei Wang
Recent advances in large language models (LLMs) have enabled intelligent tutoring systems, yet the development of LLM-based Virtual Student Agents (LVSAs) remains underexplored. Such agents are essential for teacher-facing applications, where simulating diverse learner traits can support adaptive instruction and pedagogical skill development. However, curren
Zijian Wu, Suozhi Huang, Zhejian Zhou, Huaiyuan Ying
Large Language Models (LLMs) have emerged as powerful tools in mathematical theorem proving, particularly when utilizing formal languages such as LEAN. A prevalent proof method involves the LLM prover iteratively constructing the proof tactic by tactic, typically following a best-first search scheme. However, this method often ignores the critical preference
Modified Characteristics of Hadronic Interactions in Ultra-high-energy Cosmic-ray Showers
astro-ph.HEJan Ebr, Jiri Blazek, Jakub Vicha, Tanguy Pierog
Data from multiple experiments suggest that the current interaction models used in Monte Carlo simulations do not correctly reproduce the hadronic interactions in air showers produced by ultra-high-energy cosmic rays (UHECR), in particular - but not limited to - the production of muons during the showers. We have created a large library of UHECR simulations
Jifeng Hu, Sili Huang, Li Shen, Zhejian Yang
Continual offline reinforcement learning (CORL) has shown impressive ability in diffusion-based lifelong learning systems by modeling the joint distributions of trajectories. However, most research only focuses on limited continual task settings where the tasks have the same observation and action space, which deviates from the realistic demands of training
N. N. Skryabin, Yu. A. Biriukov, M. A. Dryazgov, S. A. Fldzhyan
We present an experimental platform for linear-optical quantum information processing. Our setup utilizes multiphoton generation using a high-quality single-photon source, which is demultiplexed across multiple spatial channels, a custom-designed, programmable, low-loss photonic chip, and paired with high-efficiency single-photon detectors. We demonstrate th
Marco Cognetta, Naoaki Okazaki
Tokenization is the first step in modern neural language model pipelines where an input text is converted to a sequence of subword tokens. We introduce from first principles a finite-state transduction framework which can efficiently encode all possible tokenizations of a regular language. We then constructively show that Byte-Pair Encoding (BPE) and MaxMatc
Yiwen Pan, Wenbin Yan
In this letter, we propose a 4d mirror symmetry for the class-$\mathcal{S}$ theories which relates the representation theory of the chiral quantization of the Higgs branch and the geometry of the Coulomb branch. We study the representation theory by using the 4d/VOA correspondence, (defect) Schur indices and (flavor) modular differential equations, and match
Yunqian Cheng, Roberto Manduchi
In this paper, we present PALMS, an innovative indoor global localization and relocalization system for mobile smartphones that utilizes publicly available floor plans. Unlike most vision-based methods that require constant visual input, our system adopts a dynamic form of localization that considers a single instantaneous observation and odometry data. The
Minkwon Lee, Hyoil Kim, Changhee Joo
Federated learning (FL) is a decentralized AI mechanism suitable for a large number of devices like in smart IoT. A major challenge of FL is the non-IID dataset problem, originating from the heterogeneous data collected by FL participants, leading to performance deterioration of the trained global model. There have been various attempts to rectify non-IID da
Shot-noise limited, 10 MHz swept-source optical coherence tomography for retinal imaging
physics.opticsSacha Grelet, Alejandro Martinez Jimenez, Patrick B. Montague, Adrian Podoleanu
Akinetic swept-sources are essential for high-speed optical coherence tomography (OCT) imaging. Time-stretched supercontinuum (TSSC) lasers have proven to be efficient for multi-MHz swept-sources. However, lack of low-noise broadband lasers and of large dispersion devices in the water low-absorption band at 1060 nm have limited the biomedical applications of
Philippe Ben-Abdallah
A transverse radiative heat flux induced by the gradient of spin angular momentum of photons in non-reciprocal systems is predicted. This thermal analog of the inverse spin Hall effect is analyzed in magneto-optical networks exhibiting C4 symmetry, under the action of spatially variable external magnetic fields. This finding opens new avenues for thermal man
Sejoon Kim, Mingi Sung, Jeonghwan Lee, Hyunkuk Lim
Traditional machine translation methods typically involve training models directly on large parallel corpora, with limited emphasis on specialized terminology. However, In specialized fields such as patent, finance, or biomedical domains, terminology is crucial for translation, with many terms that needs to be translated following agreed-upon conventions. In
Enhancing SNN-based Spatio-Temporal Learning: A Benchmark Dataset and Cross-Modality Attention Model
cs.CVShibo Zhou, Bo Yang, Mengwen Yuan, Runhao Jiang
Spiking Neural Networks (SNNs), renowned for their low power consumption, brain-inspired architecture, and spatio-temporal representation capabilities, have garnered considerable attention in recent years. Similar to Artificial Neural Networks (ANNs), high-quality benchmark datasets are of great importance to the advances of SNNs. However, our analysis indic
MIK: Modified Isolation Kernel for Biological Sequence Visualization, Classification, and Clustering
cs.LGSarwan Ali, Prakash Chourasia, Haris Mansoor, Bipin koirala
The t-Distributed Stochastic Neighbor Embedding (t-SNE) has emerged as a popular dimensionality reduction technique for visualizing high-dimensional data. It computes pairwise similarities between data points by default using an RBF kernel and random initialization (in low-dimensional space), which successfully captures the overall structure but may struggle
DomainSum: A Hierarchical Benchmark for Fine-Grained Domain Shift in Abstractive Text Summarization
cs.CLHaohan Yuan, Haopeng Zhang
Most research on abstractive summarization focuses on single-domain applications, often neglecting how domain shifts between documents affect performance and the generalization ability of summarization models. To address this issue, we introduce DomainSum, a hierarchical benchmark designed to capture fine-grained domain shifts in abstractive summarization. W
Miao Yu, Shilong Wang, Guibin Zhang, Junyuan Mao
Large language models (LLMs) have empowered nodes within multi-agent networks with intelligence, showing growing applications in both academia and industry. However, how to prevent these networks from generating malicious information remains unexplored with previous research on single LLM's safety be challenging to transfer. In this paper, we focus on the sa
Manpreet Kaur, Arvind Kumar
We study the in-medium properties of kaons and antikaons in isospin asymmetric hot and dense resonance matter within the chiral SU(3) hadronic mean field model. Along with nucleons and hyperons, the interactions of $K$ and $\bar K$ mesons with all decuplet baryons ($\Delta^{++,+,0,-}, \Sigma^{*\pm,0},\Xi^{*0,-}, \Omega^{-}$) are explicitly considered in the
A Machine Learning Approach to Detect Strategic Behavior from Large-Population Observational Data Applied to Game Mode Prediction on a Team-Based Video Game
cs.GTBoshen Wang, Luis E. Ortiz
Modeling the strategic behavior of agents in a real-world multi-agent system using existing state-of-the-art computational game-theoretic tools can be a daunting task, especially when only the actions taken by the agents can be observed. Before attempting such a task, it would be useful to gain insight into whether or not agents are in fact acting strategica
Richard Dodson, Alex Williamson, Qian Gong, Pascal Elahi
The next-generation radio astronomy instruments are providing a massive increase in sensitivity and coverage, through increased stations in the array and frequency span. Two primary problems encountered when processing the resultant avalanche of data are the need for abundant storage and I/O. An example of this is the data deluge expected from the SKA Telesc
Pengcheng Shi, Shaocheng Yan, Yilin Xiao, Xinyi Liu
Correspondence-based point cloud registration (PCR) plays a key role in robotics and computer vision. However, challenges like sensor noises, object occlusions, and descriptor limitations inevitably result in numerous outliers. RANSAC family is the most popular outlier removal solution. However, the requisite iterations escalate exponentially with the outlie
Vikash Kumar Ojha, Ramkumar Radhakrishnan, Siddharth Kumar Tiwari, Mariyah Ughradar
We use phase space distributions specifically, the Wigner distribution (WD) and Husimi distribution (HD) to investigate certain information-theoretic measures as descriptors for a given system. We extensively investigate and analyze Shannon, Wehrl and Renyi entropies, its divergences, mutual information and other correlation measures within the context of th
Nishant S. Gaikwad, Lucas Heublein, Nisha L. Raichur, Tobias Feigl
Federated learning (FL) enables multiple devices to collaboratively train a global model while maintaining data on local servers. Each device trains the model on its local server and shares only the model updates (i.e., gradient weights) during the aggregation step. A significant challenge in FL is managing the feature distribution of novel and unbalanced da
Intrinsic electromagnetic damping in superconductor-ferromagnet proximity heterostructures
cond-mat.supr-conDmitriy Seleznyov, Yaroslav Turkin, Natalia Pugach, Lingling Tao
The study of the response of superconducting hybrid structures with magnetic materials to microwave irradiation is necessary for the development of effective superconducting spintronic devices. The role of the magnetic proximity effect (direct and inverse) on the electrical properties of hybrid structures is a pressing issue for its application. We theoretic
Commutativity and non-commutativity of limits in the nonlinear bending theory for prestrained microheterogeneous plates
math.APKlaus Boehnlein, Lucas Bouck, Stefan Neukamm, David Padilla-Garza
In this paper we study the derivation of nonlinear bending models for prestrained elastic plates from three-dimensional non-linear elasticity via homogenization and dimension reduction. We compare effective models obtained by either simultaneously or consecutively passing to the $\Gamma$-limits as the thickness $h\ll1$ and the size of the material microstruc
Wangjie You, Zecheng Tang, Juntao Li, Lili Yao
Large language models (LLMs) have advanced significantly due to the attention mechanism, but their quadratic complexity and linear memory demands limit their performance on long-context tasks. Recently, researchers introduced Mamba, an advanced model built upon State Space Models(SSMs) that offers linear complexity and constant memory. Although Mamba is repo
Leo Liberti, Carlile Lavor
We survey theoretical, algorithmic, and computational results at the intersection of distance geometry problems and mathematical programming, both with and without adjacencies as part of the input. While mathematical programming methods can solve large-scale distance geometry problems with adjacencies, they are severely challenged in the absence thereof.
Inverse scattering transform for the defocusing-defocusing coupled Hirota equations with non-parallel boundary conditions at infinity
nlin.SIPeng-Fei Han, Wen-Xiu Ma, Yi Zhang
The inverse scattering transform for the defocusing-defocusing coupled Hirota equations is strictly discussed with non-zero boundary conditions at infinity including non-parallel boundary conditions, specifically referring to the asymptotic polarization vectors. To address the non-analyticity encountered in some of the Jost eigenfunctions, the "adjoint" Lax
Zixuan Xu, Sibo Zheng
Using low redshift data on astrophysical reionization, we report new Lyman-$\alpha$ limit on axion-like particle (ALP) as cold dark matter in ALP mass range of $m_{a}\sim 30-1000$ eV. Compared to the Leo T and soft-X ray bound, this limit is so far the most stringent in the ALP mass range of $m_{a}\sim 375-425$ eV and complementary in the ALP mass range othe
Hyun-Kurl Jang, Jihun Kim, Hyeokjun Kweon, Kuk-Jin Yoon
Semantic Scene Completion (SSC) aims to perform geometric completion and semantic segmentation simultaneously. Despite the promising results achieved by existing studies, the inherently ill-posed nature of the task presents significant challenges in diverse driving scenarios. This paper introduces TALoS, a novel test-time adaptation approach for SSC that exc
Hongliang Lu, Xinxin Ma
Let $n,k,s$ be three integers such that $k\geq 2$ and $n\geq s\geq 1$. Let $H$ be a $k$-partite $k$-uniform hypergraph with $n$ vertices in each class. Aharoni (2017) showed that if $e(H)>(s-1)n^{k-1}$, then $H$ has a matching of size $s$. In this paper, we give a stability result for 3-partite 3-uniform hypergraphs: if $G$ is a $3$-partite $3$-uniform hyper
Regularity of Solutions for Peridynamics Equilibrium and Evolution Equations on Periodic Distributions
math.APThinh Dang, Bacim Alali, Nathan Albin
Results on the peridynamics equilibrium and evolution equations over the space of periodic vector-distributions in multi-spatial dimensions are presented. The associated operator considered is the linear state-based peridynamic operator for a homogeneous material. Results for weakly singular (integrable) as well as singular integral kernels are developed. Th
Danu Kim
Recognizing a traffic signal, determining if the signal is green or red, and figuring out the time left to cross the crosswalk are significant challenges to visually impaired people. Previous research has focused on recognizing only two traffic signals, green and red lights, using machine learning techniques. The proposed method developed a GreenEye system t
Robert Baraldi, Paul Manns
Total variation integer optimal control problems admit solutions and necessary optimality conditions via geometric variational analysis. In spite of the existence of said solutions, algorithms which solve the discretized objective suffer from high numerical cost associated with the combinatorial nature of integer programming. Hence, such methods are often li
Transforming Blood Cell Detection and Classification with Advanced Deep Learning Models: A Comparative Study
eess.IVShilpa Choudhary, Sandeep Kumar, Pammi Sri Siddhaarth, Guntu Charitasri
Efficient detection and classification of blood cells are vital for accurate diagnosis and effective treatment of blood disorders. This study utilizes a YOLOv10 model trained on Roboflow data with images resized to 640x640 pixels across varying epochs. The results show that increased training epochs significantly enhance accuracy, precision, and recall, part
Darius Feher, Abdullah Khered, Hao Zhang, Riza Batista-Navarro
In an era increasingly dominated by digital platforms, the spread of misinformation poses a significant challenge, highlighting the need for solutions capable of assessing information veracity. Our research contributes to the field of Explainable Artificial Antelligence (XAI) by developing transformer-based fact-checking models that contextualise and justify
Perturbative gradient flow coupling of the twisted Eguchi-Kawai model with the numerical stochastic perturbation theory
hep-latKen-Ichi Ishikawa, Masanori Okawa, Hironori Takei
The gradient flow scheme has emerged as a prominent nonperturbative renormalization scheme on the lattice, where flow time is introduced to define the renormalization scale. In this study we perturbatively compute the gradient flow coupling for the SU($N$) Yang-Mills theory in the large-$N$ limit in terms of the lattice bare coupling up to three-loop order.
Towards More Accurate US Presidential Election via Multi-step Reasoning with Large Language Models
cs.AIChenxiao Yu, Zhaotian Weng, Yuangang Li, Zheng Li
Can Large Language Models (LLMs) accurately predict election outcomes? While LLMs have demonstrated impressive performance in various domains, including healthcare, legal analysis, and creative tasks, their ability to forecast elections remains unknown. Election prediction poses unique challenges, such as limited voter-level data, rapidly changing political
Changmao Li, Jeffrey Flanigan
Large Language Models (LLMs) exhibit impressive results across a wide range of natural language processing (NLP) tasks, yet they can often produce factually incorrect outputs. This paper introduces a simple but effective low-latency post-correction method, \textbf{Retrieval Augmented Correction (RAC)}, aimed at enhancing the factual performance of LLMs witho
S. J. Evans, A. P. Veselov, B. Winn
A few years ago Morier-Genoud and Ovsienko introduced an interesting quantization of the real numbers as certain power series in a quantization parameter $q.$ It is known now that the golden ratio has minimal radius among all these series. We study the rational numbers having maximal radius of convergence equal to 1, which we call Kronecker fractions. We pro
Xun Jiang, Feng Li, Han Zhao, Jiahao Qiu
Large language models (LLMs) like GPTs, trained on vast datasets, have demonstrated impressive capabilities in language understanding, reasoning, and planning, achieving human-level performance in various tasks. Most studies focus on enhancing these models by training on ever-larger datasets to build more powerful foundation models. While training stronger m
Ekaterina Shemyakova, Yagmur Yilmaz
It is well known that the chain map between the de Rham and Poisson complexes on a Poisson manifold also maps the Koszul bracket of differential forms into the Schouten bracket of multivector fields. In the generalized case of a $P_\infty$-structure, where a Poisson bivector $P$ is replaced by an arbitrary even multivector obeying $[[P,P]]=0$, an analog of t
Dead-zone-free single-beam atomic magnetometer based on free-induction-decay of Rb atoms
physics.atom-phShrey Mehta, G. K. Samanta, Raghwinder Singh Grewal
Free-induction-decay (FID) magnetometers have evolved as simple magnetic sensors for sensitive detection of unknown magnetic fields. However, these magnetometers suffer from a fundamental problem known as a "dead zone," making them insensitive to certain magnetic field directions. Here, we demonstrate a simple experimental scheme for the dead-zone-free opera
Ken Furukawa, Yoshikazu Giga, Naoto Kajiwara
We consider a free boundary problem for the heat equation with a given non-negative external heat source. On the free boundary, we impose the zero Dirichlet condition and the fixed normal derivative so that heat escapes from the boundary. In various settings, we show that there exist no solutions when the initial temperature equals the fixed temperature no m
Clara Na, Ian Magnusson, Ananya Harsh Jha, Tom Sherborne
Training data compositions for Large Language Models (LLMs) can significantly affect their downstream performance. However, a thorough data ablation study exploring large sets of candidate data mixtures is typically prohibitively expensive since the full effect is seen only after training the models; this can lead practitioners to settle for sub-optimal data
SPARC: Prediction-Based Safe Control for Coupled Controllable and Uncontrollable Agents with Conformal Predictions
eess.SYShuqi Wang, Siqi Wang, Shaoyuan Li, Xiang Yin
We investigate the problem of safe control synthesis for systems operating in environments with uncontrollable agents whose dynamics are unknown but coupled with those of the controlled system. This scenario naturally arises in various applications, such as autonomous driving and human-robot collaboration, where the behavior of uncontrollable agents, like pe
Jun Zhu, Yin Xu, Dazhi He, Haoyang Li
Integrated sensing and communication (ISAC) is a very promising technology designed to provide both high rate communication capabilities and sensing capabilities. However, in Massive Multi User Multiple-Input Multiple-Output (Massive MU MIMO-ISAC) systems, the dense user access creates a serious multi-user interference (MUI) problem, leading to degradation o
Daehwan Kim, Haejun Chung, Ikbeom Jang
Deep neural networks frequently produce overconfident, miscalibrated predictions. In ordinal classification, predictions must also adhere to a unimodal and order-consistent structure, a requirement that has dominated prior work while overlooking calibration. We formalize this joint challenge as ordinal calibration for the first time and propose the Ordinal l
Jianjun Gao, Chen Cai, Ruoyu Wang, Wenyang Liu
Human-object interaction (HOI) detection has seen advancements with Vision Language Models (VLMs), but these methods often depend on extensive manual annotations. Vision Large Language Models (VLLMs) can inherently recognize and reason about interactions at the image level but are computationally heavy and not designed for instance-level HOI detection. To ov
Vansh Kharidia, Dhruvi Paprunia, Prashasti Kanikar
This paper presents LightFusionRec, a novel lightweight cross-domain recommendation system that integrates DistilBERT for textual feature extraction and FastText for genre embedding. Important issues in recommendation systems, such as data sparsity, computational efficiency, and cold start issues, are addressed in methodology. LightFusionRec uses a small amo
Khurram Yamin, Vibhhu Sharma, Ed Kennedy, Bryan Wilder
Many applications of causal inference require using treatment effects estimated on a study population to make decisions in a separate target population. We consider the challenging setting where there are covariates that are observed in the target population that were not seen in the original study. Our goal is to estimate the tightest possible bounds on het
Designing a Dataset for Convolutional Neural Networks to Predict Space Groups Consistent with Extinction Laws
cs.NEHao Wang, Jiajun Zhong, Yikun Li, Junrong Zhang
In this paper, a dataset of one-dimensional powder diffraction patterns was designed with new strategy to train Convolutional Neural Networks for predicting space groups. The diffraction pattern was calculated based on lattice parameters and Extinction Laws, instead of the traditional approach of generating it from a crystallographic database. This paper dem
Design and Optimization of a Metamaterial Absorber for Solar Energy Harvesting in the THz Frequency Range
physics.opticsNafisa Anjum, Alok Kumar Paul
This paper introduces the design and comprehensive characterization of a novel three-layer metamaterial absorber, engineered to exploit the unique optical properties of gold, vanadium dioxide, and silicon dioxide. At the core of this design, silicon dioxide serves as a robust substrate that supports an intricately structured layer of gold and a top layer of
Akshar Prabhu Desai, Ganesh Satish Mallya, Mohammad Luqman, Tejasvi Ravi
Gen-AI techniques are able to improve understanding of context and nuances in language modeling, translation between languages, handle large volumes of data, provide fast, low-latency responses and can be fine-tuned for various tasks and domains. In this manuscript, we present a comprehensive overview of the applications of Gen-AI techniques in the finance d
Yiyun He, Ke Wang, Yizhe Zhu
We derive new Hanson-Wright-type inequalities tailored to the quadratic forms of random vectors with sparse independent components. Specifically, we consider cases where the components of the random vector are sparse $\alpha$-subexponential random variables with $\alpha>0$. When $\alpha=\infty$, these inequalities can be seen as quadratic generalizations of
Jin Zhou, Hanmei Yang, Steven, Tang
Fine-tuning with Reinforcement Learning with Human Feedback (RLHF) is essential for aligning large language models (LLMs). However, RLHF often encounters significant memory challenges. This study is the first to examine memory usage in the RLHF context, exploring various memory management strategies and unveiling the reasons behind excessive memory consumpti
Richard Fang, Dylan Bowman, Daniel Kang
Recent advances in multi-modal, highly capable LLMs have enabled voice-enabled AI agents. These agents are enabling new applications, such as voice-enabled autonomous customer service. However, with all AI capabilities, these new capabilities have the potential for dual use. In this work, we show that voice-enabled AI agents can perform the actions necessary
D. C. Gunawardhana, G. H. J. Lanel, K. K. K. R. Perera, A. G. M. J. Gunaratna
This work aims to assess the molecular architectures of anti-tuberculosis drugs using both degree-based topological indices and novel distance based indices. We can represent the chemical arrangement as a graph, with atoms serving as the vertices and connections as the edges. Here, the multi bonds were considered as multi edges and included all the hydrogen
Guiwen Jiang, Chenye Qin, Kateryna Foyevtsova, Liang Si
Research on nickel-based superconductors has progressed from infinite-layer LaNiO$_2$ to finite-layer La$_{6}$Ni$_{5}$O$_{12}$, and most recently to the Ruddlesden-Popper phase La$_3$Ni$_2$O$_7$, which was found to exhibits onset of superconductivity at $\sim$80\,K under a pressure of $\sim$16\,GPa. Unlike the superconductivity mainly driven by the $d_{x^2-y
Debo Cheng, Ziqi Xu, Jiuyong Li, Lin Liu
Intervention intuition is often used in model explanation where the intervention effect of a feature on the outcome is quantified by the difference of a model prediction when the feature value is changed from the current value to the baseline value. Such a model intervention effect of a feature is inherently association. In this paper, we will study the cond
Effect of Magnetic Field on the Formation of Radiatively Inefficient Accretion Flow around Black Holes
astro-ph.HEAnish Sarkar, Mayukh Pahari
We study the effects of magnetic field in the formation of a radiatively inefficient accretion flow (RIAF) in the presence of Bremsstrahlung cooling, which facilitates the formation of a geometrically thin, optically thick accretion disk surrounded by a hot corona. We have performed axis-symmetric magnetohydrodynamic (MHD) simulations of an initial accretion
Jun Wu, Weijie Yuan, Zhiqiang Wei, Kecheng Zhang
Orthogonal time frequency space (OTFS) modulation is anticipated to be a promising candidate for supporting integrated sensing and communications (ISAC) systems, which is considered as a pivotal technique for realizing next generation wireless networks. In this paper, we develop a minimum bit error rate (BER) precoder design for an OTFS-based ISAC system. In
Hanqing Liu, Lifeng Zhou, Huanqian Yan
Large language models have drawn significant attention to the challenge of safe alignment, especially regarding jailbreak attacks that circumvent security measures to produce harmful content. To address the limitations of existing methods like GCG, which perform well in single-model attacks but lack transferability, we propose several enhancements, including
Mahdi Farrokhi Maleki, Richard Zhao
Procedural Content Generation (PCG) is defined as the automatic creation of game content using algorithms. PCG has a long history in both the game industry and the academic world. It can increase player engagement and ease the work of game designers. While recent advances in deep learning approaches in PCG have enabled researchers and practitioners to create
Chunbo Hua, Dong-Hui Xu
In recent years, there has been a surge of interest in higher-order topological phases (HOTPs) across various disciplines within the field of physics. These unique phases are characterized by their ability to harbor topological protected boundary states at lower-dimensional boundaries, a distinguishing feature that sets them apart from conventional topologic
Abdullah, Ameer Hamza, Seong Tae Kim
Medical report generation is the task of automatically writing radiology reports for chest X-ray images. Manually composing these reports is a time-consuming process that is also prone to human errors. Generating medical reports can therefore help reduce the burden on radiologists. In other words, we can promote greater clinical automation in the medical dom
Aidan Wong, He Cao, Zijing Liu, Yu Li
The increasing integration of large language models (LLMs) across various fields has heightened concerns about their potential to propagate dangerous information. This paper specifically explores the security vulnerabilities of LLMs within the field of chemistry, particularly their capacity to provide instructions for synthesizing hazardous substances. We ev
Jun Kato, Airi Mita, Keita Gobara, Akihiro Inokuchi
Graphs are useful for representing various realworld objects. However, graph neural networks (GNNs) tend to suffer from over-smoothing, where the representations of nodes of different classes become similar as the number of layers increases, leading to performance degradation. A method that does not require protracted tuning of the number of layers is needed
Can Large Language Models Invent Algorithms to Improve Themselves?: Algorithm Discovery for Recursive Self-Improvement through Reinforcement Learning
cs.CLYoichi Ishibashi, Taro Yano, Masafumi Oyamada
Large Language Models (LLMs) have achieved remarkable capabilities, yet their improvement methods remain fundamentally constrained by human design. We present Self-Developing, a framework that enables LLMs to autonomously discover, implement, and refine their own improvement algorithms. Our approach employs an iterative cycle where a seed model generates alg
Sung Gi Park, Mihnea Popa
We prove new results concerning the topology and Hodge theory of singular varieties. A common theme is that concrete conditions on the complexity of the singularities, from a number of different perspectives, are closely related to the symmetries of the Hodge-Du Bois diamond. We relate this to the theory of rational homology manifolds, and characterize these
Large Deviation Upper Bounds and Improved MSE Rates of Nonlinear SGD: Heavy-tailed Noise and Power of Symmetry
cs.LGAleksandar Armacki, Shuhua Yu, Dragana Bajovic, Dusan Jakovetic
We study large deviation upper bounds and mean-squared error (MSE) guarantees of a general framework of nonlinear stochastic gradient methods in the online setting, in the presence of heavy-tailed noise. Unlike existing works that rely on the closed form of a nonlinearity (typically clipping), our framework treats the nonlinearity in a black-box manner, allo
Yuma Kinoshita, Hitoshi Kiya
We propose a novel scene-segmentation-based exposure compensation method for multi-exposure image fusion (MEF) based tone mapping. The aim of MEF-based tone mapping is to display high dynamic range (HDR) images on devices with limited dynamic range. To achieve this, this method generates a stack of differently exposed images from an input HDR image and fuses
Hao He, Yixun Liang, Luozhou Wang, Yuanhao Cai
Recent large reconstruction models have made notable progress in generating high-quality 3D objects from single images. However, current reconstruction methods often rely on explicit camera pose estimation or fixed viewpoints, restricting their flexibility and practical applicability. We reformulate 3D reconstruction as image-to-image translation and introdu
N. T. Duy, D. T. Huong, Duong Van Loi, Phung Van Dong
We investigate a family-nonuniversal Abelian extension of hypercharge, which significantly alters the phenomenological features of the standard model. Anomaly cancellation requires that the third quark family transforms differently from the first two quark families. Additionally, it acquires that three right-handed neutrinos are presented. This model generat
Zhaonan Qu, Yongchan Kwon
Instrumental variables (IV) estimation is a fundamental method in econometrics and statistics for estimating causal effects in the presence of unobserved confounding. However, challenges such as untestable model assumptions and poor finite sample properties have undermined its reliability in practice. Viewing common issues in IV estimation as distributional
Shuzheng Si, Haozhe Zhao, Gang Chen, Yunshui Li
Aligning large language models to handle instructions with extremely long contexts has yet to be fully investigated. Previous studies have attempted to scale up the available data volume by synthesizing long instruction-following samples, as constructing such a dataset tends to be challenging for annotators. However, a lack of a well-defined strategy for ens
Thomas Kabelitz, Waseem Kamleh, Derek Leinweber
The quark mass dependence of octet baryon magnetic polarisabilities is examined at the level of individual quark-sector contributions in the uniform background-field approach of lattice QCD. The aim is to understand the direct impact of increasing the mass of a quark flavour on the magnetic polarisability and indirect or environmental effects associated with
Yuchen Chen, Weisong Sun, Chunrong Fang, Zhenpeng Chen
Language models for code (CodeLMs) have emerged as powerful tools for code-related tasks, outperforming traditional methods and standard machine learning approaches. However, these models are susceptible to security vulnerabilities, drawing increasing research attention from domains such as software engineering, artificial intelligence, and cybersecurity. De