May 2024 arXiv papers — page 74
Showing 7,301–7,400 of 20,894 papers
Less for More: Enhanced Feedback-aligned Mixed LLMs for Molecule Caption Generation and Fine-Grained NLI Evaluation
cs.CLDimitris Gkoumas, Maria Liakata
Scientific language models drive research innovation but require extensive fine-tuning on large datasets. This work enhances such models by improving their inference and evaluation capabilities with minimal or no additional training. Focusing on molecule caption generation, we explore post-training synergies between alignment fine-tuning and model merging in
Yu Shee, Anton Morgunov, Haote Li, Victor S. Batista
Traditional computer-aided synthesis planning (CASP) methods rely on iterative single-step predictions, leading to exponential search space growth that limits efficiency and scalability. We introduce a series of transformer-based models, that leverage a mixture of experts approach to directly generate multistep synthetic routes as a single string, conditiona
Ling-Qi Zhang, Zahra Kadkhodaie, Eero P. Simoncelli, David H. Brainard
We examine the problem of selecting a small set of linear measurements for reconstructing high-dimensional signals. Well-established methods for optimizing such measurements include principal component analysis (PCA), independent component analysis (ICA) and compressed sensing (CS) based on random projections, all of which rely on axis- or subspace-aligned s
Srija Chakraborty, Eleanor Stokes, Olivia Alexander
During the COVID-19 pandemic changes in human activity became widespread through official policies and organically in response to the virus's transmission, which in turn, impacted the environment and the economy. The pandemic has been described as a natural experiment that tested how social and economic disruptions impacted different components of the global
On some Analytic Inequalities for Gauss Hypergeometric Functions via Gruss Discrete Inequality
math.CAMustapha Raissouli, Mohamed Chergui
Recently, many researchers devoted their attention to study the extensions of the gamma and beta functions. In the present work, we focus on investigating some approximations for a class of Gauss hypergeometric functions by exploiting Gr\"{u}ss discrete inequality.
Nicolas Jaramillo Torres
There is an action of $\mathbb{Z}/2$ on the category of Soergel Bimodules of type $A_1 \times A_1$ induced by the nontrivial automorphism of its Dynkin diagram. We give an isotopy presentation by local generators and relations of the equivariantization of the category of Soergel Bimodules of type $A_1 \times A_1$ under this action. This is the first step in
Benchmarking the Non-flow Contributions to the Elliptic Flow Parameter ($v_{_2}$) in Proton-Proton Collisions
hep-phM. I. Abdulhamid, A. M. Hamed, E. A. Osama, M. Rateb
This manuscript reports on the elliptic flow parameter, $v_{_2}$, in proton-proton (p-p) collisions using PYTHIA8 event generator simulations. The typical Event Plane ($EP$) method $v_{_2}$($EP$) as the one used for the real data analysis has been adopted to measure the $v_{_2}$ of particles composed of different quark flavors, and produced at mid-pseudorapi
Jad Mounayer, Sebastian Rodriguez, Chady Ghnatios, Charbel Farhat
The choice of an appropriate bottleneck dimension and the application of effective regularization are both essential for Autoencoders to learn meaningful representations from unlabeled data. In this paper, we introduce a new class of deterministic autoencoders, Rank Reduction Autoencoders (RRAEs), which regularize their latent spaces by employing a truncated
Ahmad Bdeir, Johannes Burchert, Lars Schmidt-Thieme, Niels Landwehr
Hyperbolic deep learning has become a growing research direction in computer vision due to the unique properties afforded by the alternate embedding space. The negative curvature and exponentially growing distance metric provide a natural framework for capturing hierarchical relationships between datapoints and allowing for finer separability between their e
Mitigating Interference in the Knowledge Continuum through Attention-Guided Incremental Learning
cs.LGPrashant Bhat, Bharath Renjith, Elahe Arani, Bahram Zonooz
Continual learning (CL) remains a significant challenge for deep neural networks, as it is prone to forgetting previously acquired knowledge. Several approaches have been proposed in the literature, such as experience rehearsal, regularization, and parameter isolation, to address this problem. Although almost zero forgetting can be achieved in task-increment
Paul Mayer, Lorenzo Luzi, Ali Siahkoohi, Don H. Johnson
Generative models unfairly penalize data belonging to minority classes, suffer from model autophagy disorder (MADness), and learn biased estimates of the underlying distribution parameters. Our theoretical and empirical results show that training generative models with intentionally designed hypernetworks leads to models that 1) are more fair when generating
Lars Graf, Zhe Su, Giacomo Indiveri
The drive to develop artificial neural networks that efficiently utilize resources has generated significant interest in bio-inspired Spiking Neural Networks (SNNs). These networks are particularly attractive due to their potential in applications requiring low power and memory. This potential is further enhanced by the ability to perform online local learni
Annan Yu, Michael W. Mahoney, N. Benjamin Erichson
State-space models (SSMs) that utilize linear, time-invariant (LTI) systems are known for their effectiveness in learning long sequences. To achieve state-of-the-art performance, an SSM often needs a specifically designed initialization, and the training of state matrices is on a logarithmic scale with a very small learning rate. To understand these choices
Giada Pistilli, Alina Leidinger, Yacine Jernite, Atoosa Kasirzadeh
This paper introduces the "CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal impacts" dataset, designed to evaluate the social and cultural variation of Large Language Models (LLMs) across multiple languages and value-sensitive topics. We create a hand-crafted, multilingual dataset of value-laden prompts which address specific socially sensi
Numerical Simulations of 3D Ion Crystal Dynamics in a Penning Trap using the Fast Multipole Method
quant-phJohn Zaris, Wes Johnson, Athreya Shankar, John J. Bollinger
We simulate the dynamics, including laser cooling, of 3D ion crystals confined in a Penning trap using a newly developed molecular dynamics-like code. The numerical integration of the ions' equations of motion is accelerated using the fast multipole method to calculate the Coulomb interaction between ions, which allows us to efficiently study large ion cryst
Chenhui Xu, Fuxun Yu, Maoliang Li, Zihao Zheng
The past neural network design has largely focused on feature representation space dimension and its capacity scaling (e.g., width, depth), but overlooked the feature interaction space scaling. Recent advancements have shown shifted focus towards element-wise multiplication to facilitate higher-dimensional feature interaction space for better information tra
Synchrotron radiation-based tomography of an entire mouse brain with sub-micron voxels: augmenting interactive brain atlases with terabyte data
physics.med-phMattia Humbel, Christine Tanner, Marta Girona Alarcón, Georg Schulz
Synchrotron radiation-based X-ray microtomography is uniquely suited for post mortem three-dimensional visualization of organs such as the mouse brain. Tomographic imaging of the entire mouse brain with isotropic cellular resolution requires an extended field-of-view and produces datasets of multiple terabytes in size. These data must be processed and made a
Marcos Matabuena, Rahul Ghosal, Pavlo Mozharovskyi, Oscar Hernan Madrid Padilla
Depth measures have gained popularity in the statistical literature for defining level sets in complex data structures like multivariate data, functional data, and graphs. Despite their versatility, integrating depth measures into regression modeling for establishing prediction regions remains underexplored. To address this gap, we propose a novel method uti
Mahsa Golchoubian, Moojan Ghafurian, Kerstin Dautenhahn, Nasser Lashgarian Azad
Safe, socially compliant, and efficient navigation of low-speed autonomous vehicles (AVs) in pedestrian-rich environments necessitates considering pedestrians' future positions and interactions with the vehicle and others. Despite the inevitable uncertainties associated with pedestrians' predicted trajectories due to their unobserved states (e.g., intent), e
Daniel Vargas-Diaz, Jisun Kim, Sulakna Karunaratna, Maegan Reinhardt
Joint reading is a key activity for early learners, with caregiver-child interactions such as questioning and feedback playing an essential role in children's cognitive and linguistic development. However, for some parents, actively engaging children in storytelling can be challenging. To address this, we introduce TaleMate a platform designed to enhance sha
Rheeya Uppaal, Apratim Dey, Yiting He, Yiqiao Zhong
Recent alignment algorithms such as direct preference optimization (DPO) have been developed to improve the safety of large language models (LLMs) by training these models to match human behaviors exemplified by preference data. However, these methods are both computationally intensive and lacking in controllability and transparency, inhibiting their widespr
Mudit Verma, Siddhant Bhambri, Subbarao Kambhampati
The reasoning abilities of Large Language Models (LLMs) remain a topic of debate. Some methods such as ReAct-based prompting, have gained popularity for claiming to enhance sequential decision-making abilities of agentic LLMs. However, it is unclear what is the source of improvement in LLM reasoning with ReAct based prompting. In this paper we examine these
Unleashing the Power of Unlabeled Data: A Self-supervised Learning Framework for Cyber Attack Detection in Smart Grids
cs.LGHanyu Zeng, Pengfei Zhou, Xin Lou, Zhen Wei Ng
Modern power grids are undergoing significant changes driven by information and communication technologies (ICTs), and evolving into smart grids with higher efficiency and lower operation cost. Using ICTs, however, comes with an inevitable side effect that makes the power system more vulnerable to cyber attacks. In this paper, we propose a self-supervised le
Soumya Dutta, Faheem Nizar, Ahmad Amaan, Ayan Acharya
Ubiquitous applications of Deep neural networks (DNNs) in different artificial intelligence systems have led to their adoption in solving challenging visualization problems in recent years. While sophisticated DNNs offer an impressive generalization, it is imperative to comprehend the quality, confidence, robustness, and uncertainty associated with their pre
Ye Yuan, Youyuan Zhang, Can Chen, Haolun Wu
Offline model-based optimization (MBO) aims to maximize a black-box objective function using only an offline dataset of designs and scores. These tasks span various domains, such as robotics, material design, and protein and molecular engineering. A common approach involves training a surrogate model using existing designs and their corresponding scores, and
Tapas Singha, Siao-Fong Li, Murugappan Muthukumar
We study the role of active coupling on the transport properties of homogeneously charged macromolecules in an infinitely dilute solution. An enzyme becomes actively bound to a segment of the macromolecule, exerting an electrostatic force on it. Eventually, thermal fluctuations cause it to become unbound, introducing active coupling into the system. We study
Robust Generative Learning with Lipschitz-Regularized $\alpha$-Divergences Allows Minimal Assumptions on Target Distributions
stat.MLZiyu Chen, Hyemin Gu, Markos A. Katsoulakis, Luc Rey-Bellet
This paper demonstrates the robustness of Lipschitz-regularized $\alpha$-divergences as objective functionals in generative modeling, showing they enable stable learning across a wide range of target distributions with minimal assumptions. We establish that these divergences remain finite under a mild condition-that the source distribution has a finite first
Sakshi Choudhary, Sai Aparna Aketi, Kaushik Roy
Decentralized training enables learning with distributed datasets generated at different locations without relying on a central server. In realistic scenarios, the data distribution across these sparsely connected learning agents can be significantly heterogeneous, leading to local model over-fitting and poor global model generalization. Another challenge is
Md Ashfaq Salehin
In this work, an advanced deep reinforcement learning architecture is used to train neural network agents playing atari games. Given only the raw game pixels, action space, and reward information, the system can train agents to play any Atari game. At first, this system uses advanced techniques like deep Q-networks and dueling Q-networks to train efficient a
Fair Evaluation of Federated Learning Algorithms for Automated Breast Density Classification: The Results of the 2022 ACR-NCI-NVIDIA Federated Learning Challenge
eess.IVKendall Schmidt, Benjamin Bearce, Ken Chang, Laura Coombs
The correct interpretation of breast density is important in the assessment of breast cancer risk. AI has been shown capable of accurately predicting breast density, however, due to the differences in imaging characteristics across mammography systems, models built using data from one system do not generalize well to other systems. Though federated learning
Prajwal Naga, Dinesh Balivada, Sharath Chandra Nirmala, Poornoday Tiruveedi
This research paper aims to investigate the efficacy of decision trees in constructing intraday trading strategies using existing technical indicators for individual equities in the NIFTY50 index. Unlike conventional methods that rely on a fixed set of rules based on combinations of technical indicators developed by a human trader through their analysis, the
Pedro Fortuny Ayuso, Javier Ribón
In this paper we give an explicit solution to Zariski's moduli problem for plane branches. We compute (in an algorithmic way) the set of K\"{a}hler differentials of an irreducible germ of holomorphic plane curve. We show that there is a basis of this set whose main elements correspond to dicritical foliations. Indeed, we discuss several concepts of generatio
Priscylla Silva, Claudio T. Silva, Luis Gustavo Nonato
Machine learning and deep learning models are pivotal in educational contexts, particularly in predicting student success. Despite their widespread application, a significant gap persists in comprehending the factors influencing these models' predictions, especially in explainability within education. This work addresses this gap by employing nine distinct e
Leo Feng, Frederick Tung, Hossein Hajimirsadeghi, Mohamed Osama Ahmed
The advent of Transformers marked a significant breakthrough in sequence modelling, providing a highly performant architecture capable of leveraging GPU parallelism. However, Transformers are computationally expensive at inference time, limiting their applications, particularly in low-resource settings (e.g., mobile and embedded devices). Addressing this, we
Biologically Inspired Predictive Coding TCN-Transformer for Anticipatory Human-Robot Interaction in Shared Physical Spaces
cs.HCXiaoshan Zhou, Carol C. Menassa, Vineet R. Kamat
As mobile robots increasingly operate in environments shared with humans, proactively anticipating human motion rather than responding reactively is critical for preempting collisions during close-proximity navigation, while maintaining mobility efficiency and avoiding unnecessary yields. A timely and motivating engineering application is how autonomous vehi
Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao
Large language models (LLMs) are trained on a vast amount of human-written data, but data providers often remain uncredited. In response to this issue, data valuation (or data attribution), which quantifies the contribution or value of each data to the model output, has been discussed as a potential solution. Nevertheless, applying existing data valuation me
Guido De Philippis, Aria Halavati, Alessandro Pigati
In this article we prove that entire critical points $(u,\nabla)$ of the self-dual $U(1)$-Yang-Mills-Higgs functional $E_1$, with energy $$E_1(u,\nabla;B_R):=\int_{B_R}\left[|\nabla u|^2+\frac{(1-|u|^2)^2}{4}+|F_\nabla|^2\right]\leq(2\pi+\tau(n)) \omega_{n-2}R^{n-2}$$ for all $R>0$, have unique blow-down. Moreover, we show that they are two-dimensional in am
Fangzhao Zhang, Mert Pilanci
Recent developments in Parameter-Efficient Fine-Tuning (PEFT) methods for pretrained deep neural networks have captured widespread interest. In this work, we study the enhancement of current PEFT methods by incorporating the spectral information of pretrained weight matrices into the fine-tuning procedure. We investigate two spectral adaptation mechanisms, n
Divya Kothandaraman, Kihyuk Sohn, Ruben Villegas, Paul Voigtlaender
We present a method for multi-concept customization of pretrained text-to-video (T2V) models. Intuitively, the multi-concept customized video can be derived from the (non-linear) intersection of the video manifolds of the individual concepts, which is not straightforward to find. We hypothesize that sequential and controlled walking towards the intersection
Ivan Gvozdanović, Sonja Petrović
We consider the problem of constructing exact goodness-of-fit tests for discrete exponential family models. This classical problem remains practically unsolved for many types of structured or sparse data, as it rests on a computationally difficult core task: to produce a reliable sample from lattice points in a high-dimensional polytope. We translate the pro
Runlong He, Mengya Xu, Adrito Das, Danyal Z. Khan
Visual Question Answering (VQA) within the surgical domain, utilizing Large Language Models (LLMs), offers a distinct opportunity to improve intra-operative decision-making and facilitate intuitive surgeon-AI interaction. However, the development of LLMs for surgical VQA is hindered by the scarcity of diverse and extensive datasets with complex reasoning tas
Nnamdi Chikere, John McElroy, Yasemin Ozkan-Aydin
Robots are becoming increasingly essential for traversing complex environments such as disaster areas, extraterrestrial terrains, and marine environments. Yet, their potential is often limited by mobility and adaptability constraints. In nature, various animals have evolved finely tuned designs and anatomical features that enable efficient locomotion in dive
Coded Computing Meets Quantum Circuit Simulation: Coded Parallel Tensor Network Contraction Algorithm
cs.ITJin Lee, Sofia Gonzalez-Garcia, Zheng Zhang, Haewon Jeong
Parallel tensor network contraction algorithms have emerged as the pivotal benchmarks for assessing the classical limits of computation, exemplified by Google's demonstration of quantum supremacy through random circuit sampling. However, the massive parallelization of the algorithm makes it vulnerable to computer node failures. In this work, we apply coded c
Roy Allen
In a consideration set model, an individual maximizes utility among the considered alternatives. I relate a consideration set additive random utility model to classic discrete choice and the extended additive random utility model, in which utility can be $-\infty$ for infeasible alternatives. When observable utility shifters are bounded, all three models are
Andrea Serani, Matteo Diez
The rapidly evolving field of engineering design of functional surfaces necessitates sophisticated tools to manage the inherent complexity of high-dimensional design spaces. This survey paper offers a scoping review, i.e., a literature mapping synthesis borrowed from clinical medicine, delving into the field of design-space dimensionality reduction technique
DOGS: Distributed-Oriented Gaussian Splatting for Large-Scale 3D Reconstruction Via Gaussian Consensus
cs.CVYu Chen, Gim Hee Lee
The recent advances in 3D Gaussian Splatting (3DGS) show promising results on the novel view synthesis (NVS) task. With its superior rendering performance and high-fidelity rendering quality, 3DGS is excelling at its previous NeRF counterparts. The most recent 3DGS method focuses either on improving the instability of rendering efficiency or reducing the mod
Bartosz Wcisło
It is an open question whether compositional truth with the principle of propositional soundness ,,all arithmetical sentences which are propositional tautologies are true'' is conservative over its arithmetical base theory. In this article, we show that the principle of propositional soundness imposes some saturation-like properties on the underlying model,
Yigal Koifman, Ariel Barel, Alfred M. Bruckstein
This paper introduces a novel bio-mimetic approach for distributed control of robotic swarms, inspired by the collective behaviors of swarms in nature such as schools of fish and flocks of birds. The agents are assumed to have limited sensory perception, lack memory, be Identical, anonymous, and operate without interagent explicit communication. Despite thes
Yiling Xie, Xiaoming Huo
Adversarial training can achieve robustness against adversarial perturbations and has been widely used in machine learning models. This paper delivers a non-asymptotic consistency analysis of the adversarial training procedure under $\ell_\infty$-perturbation in high-dimensional linear regression. It will be shown that the associated convergence rate of pred
Daniel Grier, Hakop Pashayan, Luke Schaeffer
Given many copies of an unknown quantum state $\rho$, we consider the task of learning a classical description of its principal eigenstate. Namely, assuming that $\rho$ has an eigenstate $|\phi\rangle$ with (unknown) eigenvalue $\lambda > 1/2$, the goal is to learn a (classical shadows style) classical description of $|\phi\rangle$ which can later be used to
Aditya Agrawal, Matthew Hedlund, Blake Hechtman
eXmY is a novel data type for quantization of ML models. It supports both arbitrary bit widths and arbitrary integer and floating point formats. For example, it seamlessly supports 3, 5, 6, 7, 9 bit formats. For a specific bit width, say 7, it defines all possible formats e.g. e0m6, e1m5, e2m4, e3m3, e4m2, e5m1 and e6m0. For non-power of two bit widths e.g.
Xingtong Yu, Zhenghao Liu, Xinming Zhang, Yuan Fang
Dynamic graphs capture evolving interactions between entities, such as in social networks, online learning platforms, and crowdsourcing projects. For dynamic graph modeling, dynamic graph neural networks (DGNNs) have emerged as a mainstream technique. However, they are generally pre-trained on the link prediction task, leaving a significant gap from the obje
Aaron Brunk, Dennis Schumann
In this research, we introduce and investigate an approximation method that preserves the structural integrity of the non-isothermal Cahn-Hilliard-Navier-Stokes system. Our approach extends a previously proposed technique [1], which utilizes conforming (inf-sup stable) finite elements in space, coupled with implicit time discretization employing convex-conca
Out-of-plane magnetic phase diagram of Kitaev quantum spin liquid candidate Na2Co2TeO6
cond-mat.str-elShengzhi Zhang, Sangyun Lee, Eric Brosha, Qing Huang
We have investigated the magnetic properties and mapped out the phase diagram of the honeycomb magnet Na2Co2TeO6 with Co 3d7 in out-of-plane magnetic fields. This material has previously been proposed to show nearest-neighbor Kitaev interactions between Co spins and maybe even Kitaev quantum spin liquid behavior in high fields. At low magnetic fields, we obs
Xingtong Yu, Chang Zhou, Yuan Fang, Xinming Zhang
Given the ubiquity of graph data, it is intriguing to ask: Is it possible to train a graph foundation model on a broad range of graph data across diverse domains? A major hurdle toward this goal lies in the fact that graphs from different domains often exhibit profoundly divergent characteristics. Although there have been some initial efforts in integrating
Bharadwaj Madabhushi, Chandra Sekhar Mummidi, Sandip Kundu, Daniel Holcomb
Memory protection units (MPUs) are hardware-assisted security features that are commonly used in embedded processors such as the ARM 940T, Infineon TC1775, and Xilinx Zynq. MPUs partition the memory statically, and set individual protection attributes for each partition. MPUs typically define two protection domains: user mode and supervisor mode. Normally, t
Sylvain Kouemo Ngassom, Arghavan Moradi Dakhel, Florian Tambon, Foutse Khomh
LLM-based assistants, such as GitHub Copilot and ChatGPT, have the potential to generate code that fulfills a programming task described in a natural language description, referred to as a prompt. The widespread accessibility of these assistants enables users with diverse backgrounds to generate code and integrate it into software projects. However, studies
A Methodology to Identify Physical or Computational Experiment Conditions for Uncertainty Mitigation
cs.CEEfe Y. Yarbasi, Dimitri N. Mavris
Complex engineering systems require integration of simulation of sub-systems and calculation of metrics to drive design decisions. This paper introduces a methodology for designing computational or physical experiments for system-level uncertainty mitigation purposes. The methodology follows a previously determined problem ontology, where physical, functiona
AlabOS: A Python-based Reconfigurable Workflow Management Framework for Autonomous Laboratories
cond-mat.mtrl-sciYuxing Fei, Bernardus Rendy, Rishi Kumar, Olympia Dartsi
The recent advent of autonomous laboratories, coupled with algorithms for high-throughput screening and active learning, promises to accelerate materials discovery and innovation. As these autonomous systems grow in complexity, the demand for robust and efficient workflow management software becomes increasingly critical. In this paper, we introduce AlabOS,
Ivan Stošić, Ivan Damnjanović, Žarko Ranđelović
An expression is any mathematical formula that contains certain formal variables and operations to be executed in a specified order. In computer science, it is usually convenient to represent each expression in the form of an expression tree. Here, we consider only arithmetic expressions, i.e., those that contain only the four standard arithmetic operations:
Bharadwaj Madabhushi, Sandip Kundu, Daniel Holcomb
FPGA-based hardware accelerators are becoming increasingly popular due to their versatility, customizability, energy efficiency, constant latency, and scalability. FPGAs can be tailored to specific algorithms, enabling efficient hardware implementations that effectively leverage algorithm parallelism. This can lead to significant performance improvements ove
Some models are useful, but for how long?: A decision theoretic approach to choosing when to refit large-scale prediction models
stat.MEKentaro Hoffman, Stephen Salerno, Jeff Leek, Tyler McCormick
Large-scale prediction models using tools from artificial intelligence (AI) or machine learning (ML) are increasingly common across a variety of industries and scientific domains. Despite their effectiveness, training AI and ML tools at scale can cost tens or hundreds of thousands of dollars (or more); and even after a model is trained, substantial resources
Collective oscillations in a three-dimensional spin model with non-reciprocal interactions
cond-mat.stat-mechLaura Guislain, Eric Bertin
We study the onset of collective oscillations at low temperature in a three-dimensional spin model with non-reciprocal short-range interactions. Performing numerical simulations of the model, the presence of a continuous phase transition to global oscillations is confirmed by a finite-size scaling analysis. By systematically varying the interaction range, we
Narrative Review of Emotional Expression Support in XR: Psychophysiology of Speech-to-Text Interfaces
cs.HCSunday David Ubur, Denis Gracanin
This narrative review examines recent advancements, limitations, and research gaps in integrating emotional expression into speech-to-text (STT) interfaces within extended reality (XR) environments. Drawing from 37 peer-reviewed studies published between 2020 and 2024, we synthesized literature across multiple domains, including affective computing, psychoph
Xiang Geng, Ming Zhu, Jiahuan Li, Zhejian Lai
The scarcity of non-English data limits the development of non-English large language models (LLMs). Transforming English-centric LLMs to non-English has been identified as an effective and resource-efficient method. Previous works start from base LLMs and perform knowledge distillation (KD) with data generated by stronger LLMs, e.g. GPT-4. Compared to base
Cornelius Emde, Francesco Pinto, Thomas Lukasiewicz, Philip H. S. Torr
Since neural classifiers are known to be sensitive to adversarial perturbations that alter their accuracy, \textit{certification methods} have been developed to provide provable guarantees on the insensitivity of their predictions to such perturbations. Furthermore, in safety-critical applications, the frequentist interpretation of the confidence of a classi
Algebraic Conditions for Stability in Runge-Kutta Methods and Their Certification via Semidefinite Programming
math.NAAustin Juhl, David Shirokoff
In this work, we present approaches to rigorously certify $A$- and $A(\alpha)$-stability in Runge-Kutta methods through the solution of convex feasibility problems defined by linear matrix inequalities. We adopt two approaches. The first is based on sum-of-squares programming applied to the Runge-Kutta $E$-polynomial and is applicable to both $A$- and $A(\al
Nikolaos Chalmoukis, Georgios Nikolaidis
We study the boundedness and compactness properties of the generalized integration operator $T_{g,a}$ when it acts between distinct Hardy spaces in the unit disc of the complex plane. This operator has been introduced by the first author in connection to a theorem of Cohn about factorization of higher order derivatives of functions in Hardy spaces. We answer
François Bachoc, Nicolò Cesa-Bianchi, Tommaso Cesari, Roberto Colomboni
In online bilateral trade, a platform posts prices to incoming pairs of buyers and sellers that have private valuations for a certain good. If the price is lower than the buyers' valuation and higher than the sellers' valuation, then a trade takes place. Previous work focused on the platform perspective, with the goal of setting prices maximizing the gain fr
Olivia Di Matteo, Santiago Núñez-Corrales, Michał Stęchły, Steven P. Reinhardt
Experience from seven decades of classical computing suggests that a sustainable computer industry depends on a community of software engineers writing programs to address a wide variety of specific end-user needs, achieving both performance and utility in the process. Quantum computing is an emerging technology, and we do not yet have the insight to underst
Delyan Zhelyazov
We study the spectrum of the linearization around standing wave profiles for two quantum hydrodynamics systems with linear and nonlinear viscosity. The essential spectrum for such profiles is stable; we investigate the point spectrum using an Evans function technique. For both systems we show numerically that there exists a real unstable eigenvalue, thus pro
Felix Cherubini, Thierry Coquand, Matthias Ritter, David Wärn
Synthetic algebraic geometry is a new approach to algebraic geometry. It consists in using homotopy type theory extended with three axioms, together with the interpretation of these in a higher version of the Zariski topos, in order to do algebraic geometry internally to this topos. In this article, we will show basic properties of projective n-space $\mathb
Zhenyu Pan, Yoonsung Jeong, Xiaoda Liu, Han Liu
We propose a heterogeneous graph mamba network (HGMN) as the first exploration in leveraging the selective state space models (SSSMs) for heterogeneous graph learning. Compared with the literature, our HGMN overcomes two major challenges: (i) capturing long-range dependencies among heterogeneous nodes and (ii) adapting SSSMs to heterogeneous graph data. Our
Zhifei Yan
The chromatic number of a very dense random graph $G(n,p)$, with $p \ge 1 - n^{-c}$ for some constant $c > 0$, was first studied by Surya and Warnke, who conjectured that the typical deviation of $\chi(G(n,p))$ from its mean is of order $\sqrt{\mu_r}$, where $\mu_r$ is the expected number of independent sets of size $r$, and $r$ is maximal such that $\mu_r >
A geometrical description of non-Hermitian dynamics: speed limits in finite rank density operators
quant-phNiklas Hörnedal, Oskar A. Prośniak, Adolfo del Campo, Aurélia Chenu
Non-Hermitian dynamics in quantum systems preserves the rank of the state density operator. Using this insight, we develop a geometric framework to describe its time evolution. In particular, we identify mutually orthogonal coherent and incoherent directions and provide their physical interpretation. This understanding enables us to optimize the success rate
Matrix Denoising with Doubly Heteroscedastic Noise: Fundamental Limits and Optimal Spectral Methods
math.STYihan Zhang, Marco Mondelli
We study the matrix denoising problem of estimating the singular vectors of a rank-$1$ signal corrupted by noise with both column and row correlations. Existing works are either unable to pinpoint the exact asymptotic estimation error or, when they do so, the resulting approaches (e.g., based on whitening or singular value shrinkage) remain vastly suboptimal
François Bachoc, Tommaso Cesari, Roberto Colomboni
We study the role of contextual information in the online learning problem of brokerage between traders. In this sequential problem, at each time step, two traders arrive with secret valuations about an asset they wish to trade. The learner (a broker) suggests a trading (or brokerage) price based on contextual data about the asset and the market conditions.
Wei Li, Hehe Fan, Yongkang Wong, Mohan Kankanhalli
Recent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability of substantial web video-text data. This difficulty primarily arises from the inherent complexity of videos and the inefficient language supervision in recent web-collected video-t
Jiali Cui, Tian Han
This work studies the learning problem of the energy-based prior model and the multi-layer generator model. The multi-layer generator model, which contains multiple layers of latent variables organized in a top-down hierarchical structure, typically assumes the Gaussian prior model. Such a prior model can be limited in modelling expressivity, which results i
Vu Hoang
In this paper, we consider a theory defined by an energy-momentum tensor depending on a set of general fields, including the space-time metric. We prove that if the theory is causal, bounded and transforms appropriately under diffeomorphism, it will depend only on the local values of the independent fields and their covariant derivatives up to a finite order
Ludovic D. C. Jaubert, Yasir Iqbal, Harald O. Jeschke
The chromium spinels MgCr2O4 and ZnCr2O4 are prime examples of the highly frustrated pyrochlore lattice antiferromagnet. Experiment has carefully established that both materials, upon cooling, distort to lower symmetry and order magnetically. We study the nature of this process by a combination of density-functional-theory based energy mapping and classical
Just rephrase it! Uncertainty estimation in closed-source language models via multiple rephrased queries
cs.CLAdam Yang, Chen Chen, Konstantinos Pitas
State-of-the-art large language models are sometimes distributed as open-source software but are also increasingly provided as a closed-source service. These closed-source large-language models typically see the widest usage by the public, however, they often do not provide an estimate of their uncertainty when responding to queries. As even the best models
O. L. Dors, M. V. Cardaci, G. F. Hägele, M. Valerdi
We derive the nitrogen and oxygen abundances in the Narrow Line Regions (NLRs) of a sample of 38 local ($z \: < \: 0.4$) Seyfert~2 nuclei. For that, we consider narrow optical emission line intensities and direct estimates of the electron temperatures ($T_{\rm e}$-method). We find nitrogen abundances in the range $7.6 \: < \: \rm 12+log(N/H) \: < \: 8.6$ (me
Calibration of stochastic, agent-based neuron growth models with Approximate Bayesian Computation
cs.CETobias Duswald, Lukas Breitwieser, Thomas Thorne, Barbara Wohlmuth
Understanding how genetically encoded rules drive and guide complex neuronal growth processes is essential to comprehending the brain's architecture, and agent-based models (ABMs) offer a powerful simulation approach to further develop this understanding. However, accurately calibrating these models remains a challenge. Here, we present a novel application o
ECG-TEM: Time-based sub-Nyquist sampling for ECG signal reconstruction and Hardware Prototype
eess.SPHila Naaman, Daniel Bilik, Shlomi Savariego, Moshe Namer
Portable heart rate monitoring (HRM) systems based on electrocardiograms (ECGs) have become increasingly crucial for preventing lifestyle diseases. For such portable systems, minimizing power consumption and sampling rate is critical due to the substantial data generated during long-term ECG monitoring. The variable pulse-width finite rate of innovation (VPW
ST-Gait++: Leveraging spatio-temporal convolutions for gait-based emotion recognition on videos
cs.CVMaria Luísa Lima, Willams de Lima Costa, Estefania Talavera Martinez, Veronica Teichrieb
Emotion recognition is relevant for human behaviour understanding, where facial expression and speech recognition have been widely explored by the computer vision community. Literature in the field of behavioural psychology indicates that gait, described as the way a person walks, is an additional indicator of emotions. In this work, we propose a deep framew
Shoshana Abramovich
In this paper we prove results on the difference between a normalized Jensen functional and the sum of other normalized Jensen functionals for convex function.
Yiran Qiao, Xiang Ao, Yang Liu, Jiarong Xu
Recent prevailing works on graph machine learning typically follow a similar methodology that involves designing advanced variants of graph neural networks (GNNs) to maintain the superior performance of GNNs on different graphs. In this paper, we aim to streamline the GNN design process and leverage the advantages of Large Language Models (LLMs) to improve t
Rui Sun, Haoran Duan, Jiahua Dong, Varun Ojha
We introduce a rehearsal-free federated domain incremental learning framework, RefFiL, based on a global prompt-sharing paradigm to alleviate catastrophic forgetting challenges in federated domain-incremental learning, where unseen domains are continually learned. Typical methods for mitigating forgetting, such as the use of additional datasets and the reten
Nam Phuong Tran, The Anh Ta, Debmalya Mandal, Long Tran-Thanh
High-dimensional linear bandits with low-dimensional structure have received considerable attention in recent studies due to their practical significance. The most common structure in the literature is sparsity. However, it may not be available in practice. Symmetry, where the reward is invariant under certain groups of transformations on the set of arms, is
Alejandro Gomez Cadavid, Archismita Dalal, Anton Simen, Enrique Solano
We introduce a method for solving combinatorial optimization problems on digital quantum computers, where we incorporate auxiliary counterdiabatic (CD) terms into the adiabatic Hamiltonian, while integrating bias terms derived from an iterative digitized counterdiabatic quantum algorithm. We call this protocol bias-field digitized counterdiabatic quantum opt
Jane Ivy Coons, Heather A. Harrington, Niharika Chakrabarty Paul
We investigate the geometry of a family of log-linear statistical models called quasi-independence models. The toric fiber product is useful for understanding the geometry of parameter inference in these models because the maximum likelihood degree is multiplicative under the TFP. We define the coordinate toric fiber product, or cTFP, and give necessary and
Maria Koshkina, James H. Elder
Jersey number recognition is an important task in sports video analysis, partly due to its importance for long-term player tracking. It can be viewed as a variant of scene text recognition. However, there is a lack of published attempts to apply scene text recognition models on jersey number data. Here we introduce a novel public jersey number recognition da
K. L. Yeo, S. K. Solanki, N. A. Krivova
We compared magnetograms from the KPVT/SPM, SoHO/MDI, SOLIS/VSM, and SDO/HMI with the aim of probing the effect on measured solar magnetism of the variation in instrument response with time, magnetogram signal level, and position on the solar disc. Taking near-simultaneous observations from the various instruments, we examined the surface coverage by magneti
Xiaozhou Feng, Nadezhda Fishchenko, Sarang Gopalakrishnan, Matteo Ippoliti
The dynamics of monitored systems can exhibit a measurement-induced phase transition (MIPT) between entangling and disentangling phases, tuned by the measurement rate. When the dynamics obeys a continuous symmetry, the entangling phase further splits into a fuzzy phase and a sharp phase based on the scaling of fluctuations of the symmetry charge. While the s
Igor Grzelec, Tomáš Madaras, Alfréd Onderko, Roman Soták
A graph/multigraph $G$ is locally irregular if endvertices of every its edge possess different degrees. The locally irregular edge coloring of $G$ is its edge coloring with the property that every color induces a locally irregular sub(multi)graph of $G$; if such a coloring of $G$ exists, the minimum number of colors to color $G$ in this way is the locally ir
Hansveer Singh, Ewan McCulloch, Sarang Gopalakrishnan, Romain Vasseur
We construct an ensemble of two-dimensional nonintegrable quantum circuits that are chaotic but have a conserved particle current, and thus a finite Drude weight. The long-wavelength hydrodynamics of such systems is given by the incompressible Navier-Stokes equations. By analyzing circuit-to-circuit fluctuations in the ensemble we argue that these are neglig
Andrew G. Sullivan, Roger W. Romani
PSR J2215+5135 (J2215) is a `redback' spider pulsar, where the intrabinary shock (IBS) wraps around the pulsar rather than the stellar-mass companion. Spider orbital light curves are modulated, dominated by their binary companion thermal emission in the optical bands and by IBS synchrotron emission in the X-rays. We report on new XMM-Newton X-ray and U-band
Dingling Yao, Caroline Muller, Francesco Locatello
Causal representation learning promises to extend causal models to hidden causal variables from raw entangled measurements. However, most progress has focused on proving identifiability results in different settings, and we are not aware of any successful real-world application. At the same time, the field of dynamical systems benefited from deep learning an
Multi-Zone Modeling of Black Hole Accretion and Feedback in 3D GRMHD: Bridging Vast Spatial and Temporal Scales
astro-ph.HEHyerin Cho, Ben S. Prather, Kung-Yi Su, Ramesh Narayan
Simulating accretion and feedback from the horizon scale of supermassive black holes (SMBHs) out to galactic scales is challenging because of the vast range of scales involved. Elaborating on H. Cho et al., we describe and test a "multizone" technique, which is designed to tackle this difficult problem in three-dimensional general relativistic magnetohydrody