May 2023 arXiv papers — page 31
Showing 3,001–3,100 of 19,695 papers
Jiayu Chen, Zhekai Wang, Vaneet Aggarwal
Imperfect Information Games (IIGs) offer robust models for scenarios where decision-makers face uncertainty or lack complete information. Counterfactual Regret Minimization (CFR) has been one of the most successful family of algorithms for tackling IIGs. The integration of skill-based strategy learning with CFR could potentially mirror more human-like decisi
Yifan Zhang, Zhiquan Tan, Jingqin Yang, Weiran Huang
The maximum entropy encoding framework provides a unified perspective for many non-contrastive learning methods like SimSiam, Barlow Twins, and MEC. Inspired by this framework, we introduce Matrix-SSL, a novel approach that leverages matrix information theory to interpret the maximum entropy encoding loss as matrix uniformity loss. Furthermore, Matrix-SSL en
Tianjian Li, Kenton Murray
Zero-shot cross-lingual transfer is when a multilingual model is trained to perform a task in one language and then is applied to another language. Although the zero-shot cross-lingual transfer approach has achieved success in various classification tasks, its performance on natural language generation tasks falls short in quality and sometimes outputs an in
Soham Mandal, Paul C. Duffell, Abigail Polin, Dan Milisavljevic
We develop a suite of 3D hydrodynamic models of supernova remnants (SNRs) expanding against the circumstellar medium (CSM). We study the Rayleigh-Taylor Instability (RTI) forming at the expansion interface by calculating an angular power spectrum for each of these models. The power spectra of young SNRs is seen to exhibit a dominant angular mode, which is a
Benjamin Grimmer, Danlin Li
We consider (stochastic) subgradient methods for strongly convex but potentially nonsmooth non-Lipschitz optimization. We provide new equivalent dual descriptions (in the style of dual averaging) for the classic subgradient method, the proximal subgradient method, and the switching subgradient method. These equivalences enable $O(1/T)$ convergence guarantees
Zi-Ang Hu, Bo Fu, Xiao Li, Shun-Qing Shen
Discrete time crystal is a class of nonequilibrium quantum systems exhibiting subharmonic responses to external periodic driving. Here we propose a class of discrete time crystals enforced by nonsymmorphic dynamical symmetry. We start with a system with nonsymmorphic dynamical symmetry, in which the instantaneous eigenstates become M\"obius twisted, hence do
Flávio G. C. Rocha, Gabriel M. F. de Almeida, Kleber V. Cardoso, Cristiano B. Both
In this article, we propose a novel formulation for the resource allocation problem of a sliced and disaggregated Radio Access Network (RAN) and its transport network. Our proposal assures an end-to-end delay bound for the Ultra-Reliable and Low-Latency Communication (URLLC) use case while jointly considering the number of admitted users, the transmission ra
Aline J. O. Andrade, Bruno L. M. Ferreira, Liudmila Sabinina
Let A and A' be two alternative *-algebras with identities 1_A and 1_A', respectively, and e_1 and e_2 = 1_A - e_1 nontrivial symmetric idempotents in A. In this paper we study the characterization of multiplicative *-Jordan-type maps on alternative algebras.
Joaquim Dias Garcia, Guilherme Bodin, Alexandre Street
In this technical report, we compare multiple reformulation techniques and solvers that can be used with the Julia package BilevelJuMP. We focus on the special case of Hyperparameter Tuning for Support Vector Regression. We describe a bilevel model for the problem in question. Then we present code for generating data and models that solve the problem. Finall
Michael Feffer, Hoda Heidari, Zachary C. Lipton
With Artificial Intelligence systems increasingly applied in consequential domains, researchers have begun to ask how these systems ought to act in ethically charged situations where even humans lack consensus. In the Moral Machine project, researchers crowdsourced answers to "Trolley Problems" concerning autonomous vehicles. Subsequently, Noothigattu et al.
Radar Enlighten the Dark: Enhancing Low-Visibility Perception for Automated Vehicles with Camera-Radar Fusion
cs.CVCan Cui, Yunsheng Ma, Juanwu Lu, Ziran Wang
Sensor fusion is a crucial augmentation technique for improving the accuracy and reliability of perception systems for automated vehicles under diverse driving conditions. However, adverse weather and low-light conditions remain challenging, where sensor performance degrades significantly, exposing vehicle safety to potential risks. Advanced sensors such as
Allison Sullivan
Finite model finders give users the ability to specify properties of a system in mathematical logic and then automatically find concrete examples, called solutions, that satisfy the properties. These solutions are often viewed as a key benefit of model finders, as they create an exploratory environment for developers to engage with their model. In practice,
Dihua Jiang, Zhaolin Li, Guodong Xi
We prove the uniqueness of the Ginzburg-Rallis models over $p$-adic local fields of characteristic zero, which completes the local uniqueness problem for the Ginzburg-Rallis models starting from the work of C.-F. Nien in \cite{MR2709083} that proves the non-split case, and the work of D. Jiang, B. Sun and C. Zhu in \cite{MR2763736} that proves the general ca
Qin Wu, Zhen-Yin Zhao, F. Y. Wang
Recently, remarkable anti-glitch and glitch accompanied by bright radio bursts of the Galactic magnetar SGR J1935+2154 were discovered. These two infrequent temporal coincidences between the glitch/anti-glitch and the fast radio burst (FRB)-like bursts reveal their physical connection of them. Here we propose that the anti-glitch/glitch and FRB-like bursts c
Shuochuan Meng, Mohammad Hesam Soleimani-Babakamali, Ertugrul Taciroglu
Roof type is one of the most critical building characteristics for wind vulnerability modeling. It is also the most frequently missing building feature from publicly available databases. An automatic roof classification framework is developed herein to generate high-resolution roof-type data using machine learning. A Convolutional Neural Network (CNN) was tr
Zezhen Sun
In this paper we introduce two $1/\kappa^{n}$-type ($n\ge1$) curvature flows for closed convex planar curves. Along the flows the length of the curve is decreasing while the enclosed area is increasing. And finally, the evolving curves converge smoothly to a finite circle if they do not develop singularity during the evolution process.
Super-Resolution of License Plate Images Using Attention Modules and Sub-Pixel Convolution Layers
cs.CVValfride Nascimento, Rayson Laroca, Jorge de A. Lambert, William Robson Schwartz
Recent years have seen significant developments in the field of License Plate Recognition (LPR) through the integration of deep learning techniques and the increasing availability of training data. Nevertheless, reconstructing license plates (LPs) from low-resolution (LR) surveillance footage remains challenging. To address this issue, we introduce a Single-
Yuhui Zhang, Michihiro Yasunaga, Zhengping Zhou, Jeff Z. HaoChen
Language models have been shown to exhibit positive scaling, where performance improves as models are scaled up in terms of size, compute, or data. In this work, we introduce NeQA, a dataset consisting of questions with negation in which language models do not exhibit straightforward positive scaling. We show that this task can exhibit inverse scaling, U-sha
Igor Nunes, Mike Heddes, Pere Vergés, Danny Abraham
Metrics for set similarity are a core aspect of several data mining tasks. To remove duplicate results in a Web search, for example, a common approach looks at the Jaccard index between all pairs of pages. In social network analysis, a much-celebrated metric is the Adamic-Adar index, widely used to compare node neighborhood sets in the important problem of p
Muhammad Inam
We introduce the notion of a subgraph generated by an $R$-word $r$ of the Schützenberger graph of a positive word $w$, $SΓ(w)$, where $w$ contains $r$ as its subword. We show that the word problem for a finitely presented Adian inverse semigroup $Inv\langle X|R \rangle$ is decidable if the subgraphs of $SΓ(t)$, for all $t\in X^+$, generated by all the $R$-wo
Hao Zhu, Yong-Yao Li, Wen-Kai Bai, Yan-Mei Yu
We predict a scheme for the creation of isotropic three-dimensional droplets in Rydbeg-dressed Bose gases, which contain both repulsive contact interactions and attractive van der Waals interactions causing the quantum fluctuation effect non-negligible. We present detailed beyond mean-field calculations with Lee-Huang-Yang correction and demonstrate the exis
Habtom Kahsay Gidey, Peter Hillmann, Andreas Karcher, Alois Knoll
Software bots operating in multiple virtual digital platforms must understand the platforms' affordances and behave like human users. Platform affordances or features differ from one application platform to another or through a life cycle, requiring such bots to be adaptable. Moreover, bots in such platforms could cooperate with humans or other software agen
Minghao Xia, Sijie Gao
Tolman proposed that the proper temper $T$ of a static self-gravitating fluid in thermodynamic equilibrium satisfies the relation $\chi T=constant$, where $\chi$ is the redshift factor of the spacetime. The Tolman law has been proven for radiation in stationary spacetimes and for perfect fluids in stationary, asymototically flat and axisymmetric spacetimes.
Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning Performance
cs.CLYao Fu, Litu Ou, Mingyu Chen, Yuhao Wan
As large language models (LLMs) are continuously being developed, their evaluation becomes increasingly important yet challenging. This work proposes Chain-of-Thought Hub, an open-source evaluation suite on the multi-step reasoning capabilities of large language models. We are interested in this setting for two reasons: (1) from the behavior of GPT and PaLM
Elahe Rahimian, Golara Javadi, Frederick Tung, Gabriel Oliveira
Multi-task networks rely on effective parameter sharing to achieve robust generalization across tasks. In this paper, we present a novel parameter sharing method for multi-task learning that conditions parameter sharing on both the task and the intermediate feature representations at inference time. In contrast to traditional parameter sharing approaches, wh
Michael Levit, Sarangarajan Parthasarathy, Cem Aksoylar, Mohammad Sadegh Rasooli
We propose an adaptation method for factorized neural transducers (FNT) with external language models. We demonstrate that both neural and n-gram external LMs add significantly more value when linearly interpolated with predictor output compared to shallow fusion, thus confirming that FNT forces the predictor to act like regular language models. Further, we
Shantanu Ghosh, Ke Yu, Kayhan Batmanghelich
Building generalizable AI models is one of the primary challenges in the healthcare domain. While radiologists rely on generalizable descriptive rules of abnormality, Neural Network (NN) models suffer even with a slight shift in input distribution (e.g., scanner type). Fine-tuning a model to transfer knowledge from one domain to another requires a significan
Haiyan Li, Ilia Ponomarenko, Peter Zeman
Let $m$ be a positive integer, $X$ a graph with vertex set $\Omega$, and ${\rm WL}_m(X)$ the coloring of the Cartesian $m$-power $\Omega^m$, obtained by the $m$-dimensional Weisfeiler-Leman algorithm. The ${\rm WL}$-dimension of the graph $X$ is defined to be the smallest $m$ for which the coloring ${\rm WL}_m(X)$ determines $X$ up to isomorphism. It is know
Stability-penalty-adaptive follow-the-regularized-leader: Sparsity, game-dependency, and best-of-both-worlds
cs.LGTaira Tsuchiya, Shinji Ito, Junya Honda
Adaptivity to the difficulties of a problem is a key property in sequential decision-making problems to broaden the applicability of algorithms. Follow-the-regularized-leader (FTRL) has recently emerged as one of the most promising approaches for obtaining various types of adaptivity in bandit problems. Aiming to further generalize this adaptivity, we develo
Exploiting Large Neuroimaging Datasets to Create Connectome-Constrained Approaches for more Robust, Efficient, and Adaptable Artificial Intelligence
cs.NEErik C. Johnson, Brian S. Robinson, Gautam K. Vallabha, Justin Joyce
Despite the progress in deep learning networks, efficient learning at the edge (enabling adaptable, low-complexity machine learning solutions) remains a critical need for defense and commercial applications. We envision a pipeline to utilize large neuroimaging datasets, including maps of the brain which capture neuron and synapse connectivity, to improve mac
Dimitris Bertsimas, Vassilis Digalakis
Owing to their inherently interpretable structure, decision trees are commonly used in applications where interpretability is essential. Recent work has focused on improving various aspects of decision trees, including their predictive power and robustness; however, their instability, albeit well-documented, has been addressed to a lesser extent. In this pap
Youngeun Kim, Yuhang Li, Abhishek Moitra, Ruokai Yin
Spiking Neural Networks (SNNs) have gained increasing attention as energy-efficient neural networks owing to their binary and asynchronous computation. However, their non-linear activation, that is Leaky-Integrate-and-Fire (LIF) neuron, requires additional memory to store a membrane voltage to capture the temporal dynamics of spikes. Although the required me
Optimizing Representation in Redistricting: Dual Bounds for Partitioning Problems with Non-Convex Objectives
math.OCJamie Fravel, Robert Hildebrand, Nicholas Goedert, Laurel Travis
We investigate optimization models for the purpose of computational redistricting. Our focus is on nonconvex objectives for estimating expected Black Representatives and Political Representation. The objectives are a composition of a ratio of variables and a normal distribution's cumulative distribution function (or ``probit curve"). We extend the work of Va
Chinmaya Kausik, Kashvi Srivastava, Rishi Sonthalia
Despite the importance of denoising in modern machine learning and ample empirical work on supervised denoising, its theoretical understanding is still relatively scarce. One concern about studying supervised denoising is that one might not always have noiseless training data from the test distribution. It is more reasonable to have access to noiseless train
J. Freedberg, W. Joe Meese, J. He, D. L. Schlagel
The memory effect in a single crystal spin glass ($\mathrm{Cu}_{0.92}\mathrm{Mn}_{0.08}$) has been measured using \freq ac susceptibility techniques over a temperature range of $0.4 - 0.7 \, T_g$ and a model of the memory effect has been developed. A double-waiting-time protocol is carried out where the spin glass is first allowed to age at a temperature bel
Alon Harell, Yalda Foroutan, Nilesh Ahuja, Parual Datta
Recent years have seen a tremendous growth in both the capability and popularity of automatic machine analysis of images and video. As a result, a growing need for efficient compression methods optimized for machine vision, rather than human vision, has emerged. To meet this growing demand, several methods have been developed for image and video coding for m
A boundary integral equation method for the complete electrode model in electrical impedance tomography with tests on experimental data
math.NATeemu Tyni, Adam R Stinchcombe, Spyros Alexakis
We develop a boundary integral equation-based numerical method to solve for the electrostatic potential in two dimensions, inside a medium with piecewise constant conductivity, where the boundary condition is given by the complete electrode model (CEM). The CEM is seen as the most accurate model of the physical setting where electrodes are placed on the surf
Arianna Cecco
We explore functors between operator space categories, some properties of these functors, and establish relations between objects in these categories and their images under these functors, in particular regarding injectivity and injective envelopes. We also compare the purely categorical definition of injectivity with the `standard' operator theoretical defi
Yago Antolín, Islam Foniqi
We prove a Tits alternative theorem for subgroups of finitely generated even Artin groups of FC type (EAFC groups), stating that there exists a finite index subgroup such that every subgroup of it is either finitely generated abelian, or maps onto a non-abelian free group. Parabolic subgroups play a key role, and we show that parabolic subgroups of EAFC grou
Convex Risk Bounded Continuous-Time Trajectory Planning and Tube Design in Uncertain Nonconvex Environments
cs.AIAshkan Jasour, Weiqiao Han, Brian Williams
In this paper, we address the trajectory planning problem in uncertain nonconvex static and dynamic environments that contain obstacles with probabilistic location, size, and geometry. To address this problem, we provide a risk bounded trajectory planning method that looks for continuous-time trajectories with guaranteed bounded risk over the planning time h
Beatrice Andreolli, Karlheinz Gröchenig
We introduce a new concept of variable bandwidth that is based on the truncation of Wilson expansions. For this model we derive both (nonuniform) sampling theorems, the complete reconstruction of $f$ from its samples, and necessary density conditions for sampling.
Fourier-DeepONet: Fourier-enhanced deep operator networks for full waveform inversion with improved accuracy, generalizability, and robustness
cs.LGMin Zhu, Shihang Feng, Youzuo Lin, Lu Lu
Full waveform inversion (FWI) infers the subsurface structure information from seismic waveform data by solving a non-convex optimization problem. Data-driven FWI has been increasingly studied with various neural network architectures to improve accuracy and computational efficiency. Nevertheless, the applicability of pre-trained neural networks is severely
Sushovan Majhi
For a closed Riemannian manifold $\mathcal{M}$ and a metric space $S$ with a small Gromov$\unicode{x2013}$Hausdorff distance to it, Latschev's theorem guarantees the existence of a sufficiently small scale $\beta>0$ at which the Vietoris$\unicode{x2013}$Rips complex of $S$ is homotopy equivalent to $\mathcal{M}$. Despite being regarded as a stepping stone to
Michael Zurel, Cihan Okay, Robert Raussendorf
A recently introduced classical simulation method for universal quantum computation with magic states operates by repeated sampling from probability functions [M. Zurel et al. PRL 260404 (2020)]. This method is closely related to sampling algorithms based on Wigner functions, with the important distinction that Wigner functions can take negative values obstr
Neil A. Ernst, Martin P. Robillard
Documentation is an important mechanism for disseminating software architecture knowledge. Software project teams can employ vastly different formats for documenting software architecture, from unstructured narratives to standardized documents. We explored to what extent this documentation format may matter to newcomers joining a software project and attempt
Andrea Gabrielli, Valentina Macchiati, Diego Garlaschelli
The structure of many financial networks is protected by privacy and has to be inferred from aggregate observables. Here we consider one of the most successful network reconstruction methods, producing random graphs with desired link density and where the observed constraints (related to the market size of each node) are replicated as averages over the graph
Tianchun Wang, Farzaneh Mirzazadeh, Xiang Zhang, Jie Chen
Graph convolutional networks (GCNs) are \emph{discriminative models} that directly model the class posterior $p(y|\mathbf{x})$ for semi-supervised classification of graph data. While being effective, as a representation learning approach, the node representations extracted from a GCN often miss useful information for effective clustering, because the objecti
Aakash Lahoti, Spandan Senapati, Ketan Rajawat, Alec Koppel
The problem of minimizing the sum of $n$ functions in $d$ dimensions is ubiquitous in machine learning and statistics. In many applications where the number of observations $n$ is large, it is necessary to use incremental or stochastic methods, as their per-iteration cost is independent of $n$. Of these, Quasi-Newton (QN) methods strike a balance between the
Sushma Kumari, Vladimir G. Pestov
We continue to investigate the $k$ nearest neighbour ($k$-NN) learning rule in complete separable metric spaces. Thanks to the results of C\'erou and Guyader (2006) and Preiss (1983), this rule is known to be universally consistent in every such metric space that is sigma-finite dimensional in the sense of Nagata. Here we show that the rule is strongly unive
Fani Dosopoulou
We consider the intermediate mass-ratio inspiral of a stellar-mass compact object with an intermediate-mass black hole that is surrounded by a dark matter density spike. The interaction of the inspiraling black hole with the dark matter particles in the spike leads to dynamical friction. This can alter the dynamics of the black hole binary, leaving an imprin
Duong Minh Le, Ruohao Guo, Wei Xu, Alan Ritter
In this paper, we study the task of instructional dialogue and focus on the cooking domain. Analyzing the generated output of the GPT-J model, we reveal that the primary challenge for a recipe-grounded dialog system is how to provide the instructions in the correct order. We hypothesize that this is due to the model's lack of understanding of user intent and
Archetypal solution spaces for clustering gene expression datasets in identification of cancer subtypes
physics.bio-phYuchen Wu, Luke Dicks, David J. Wales
Gene expression profiles are essential in identifying different cancer phenotypes. Clustering gene expression datasets can provide accurate identification of cancerous cell lines, but this task is challenging due to the small sample size and high dimensionality. Using the $K$-means clustering algorithm we determine the organisation of the solution space for
One-Parameter Meromorphic Solution of the Degenerate Third Painlev\'{e} Equation with Formal Monodromy Parameter $a=\pm i/2$ Vanishing at the Origin
math.CAA. V. Kitaev, A. Vartanian
We prove that there exists a one-parameter meromorphic solution $u(\tau)$ vanishing at $\tau=0$ of the degenerate third Painlev\'e equation, \begin{equation*} u^{\prime \prime}(\tau) \! = \! \frac{(u^{\prime}(\tau))^{2}}{u(\tau)} \! - \! \frac{u^{\prime}(\tau)}{\tau} \! + \! \frac{1}{\tau} \! \left(-8 \varepsilon (u(\tau))^{2} \! + \! 2ab \right) \! + \! \fr
Chang Deng, Kevin Bello, Bryon Aragam, Pradeep Ravikumar
Recently, an intriguing class of non-convex optimization problems has emerged in the context of learning directed acyclic graphs (DAGs). These problems involve minimizing a given loss or score function, subject to a non-convex continuous constraint that penalizes the presence of cycles in a graph. In this work, we delve into the optimization challenges assoc
Differentiability of the effective Lagrangian for Hamilton-Jacobi-Bellman equations in dynamic random environments
math.PRYuri Bakhtin, Douglas Dow
We prove differentiability of the effective Lagrangian for continuous time multidimensional directed variational problems in random dynamic environments with positive dependence range in time. This implies that limiting fundamental solutions in the associated homogenization problems for HJB equations are classical.
Enia Xhakaj, Alexie Leauthaud, Johannes Lange, Elisabeth Krause
We propose that observations of super-massive galaxies contain cosmological constraining power similar to conventional cluster cosmology, and we provide promising indications that the associated systematic errors are comparably easier to control. We consider a fiducial spectroscopic and stellar mass complete sample of galaxies drawn from the Dark Energy Spec
Local Convergence of Gradient Methods for Min-Max Games: Partial Curvature Generically Suffices
math.OCGuillaume Wang, Lénaïc Chizat
We study the convergence to local Nash equilibria of gradient methods for two-player zero-sum differentiable games. It is well-known that such dynamics converge locally when $S \succ 0$ and may diverge when $S=0$, where $S\succeq 0$ is the symmetric part of the Jacobian at equilibrium that accounts for the "potential" component of the game. We show that thes
How to verify the precision of density-functional-theory implementations via reproducible and universal workflows
cond-mat.mtrl-sciEmanuele Bosoni, Louis Beal, Marnik Bercx, Peter Blaha
In the past decades many density-functional theory methods and codes adopting periodic boundary conditions have been developed and are now extensively used in condensed matter physics and materials science research. Only in 2016, however, their precision (i.e., to which extent properties computed with different codes agree among each other) was systematicall
Ruoxi Sun, Sercan Ö. Arik, Alex Muzio, Lesly Miculicich
Text-to-SQL, the process of translating natural language into Structured Query Language (SQL), represents a transformative application of large language models (LLMs), potentially revolutionizing how humans interact with data. This paper introduces the SQL-PaLM framework, a comprehensive solution for understanding and enhancing Text-to-SQL using LLMs, using
Sadhana Kumaravel, Tahira Naseem, Ramon Fernandez Astudillo, Radu Florian
The sliding window approach provides an elegant way to handle contexts of sizes larger than the Transformer's input window, for tasks like language modeling. Here we extend this approach to the sequence-to-sequence task of document parsing. For this, we exploit recent progress in transition-based parsing to implement a parser with synchronous sliding windows
Sigurd Angenent, Panagiota Daskalopoulos, Natasa Sesum
There is an extensive and growing body of work analyzing convex ancient solutions to Mean Curvature Flow (MCF), or equivalently of Rescaled Mean Curvature Flow (RMCF). The goal of this paper is to complement the existing literature, which analyzes ancient solutions one at a time, by considering the space X of all convex hypersurfaces M, regard RMCF as a semi
Robust Lane Detection through Self Pre-training with Masked Sequential Autoencoders and Fine-tuning with Customized PolyLoss
cs.CVRuohan Li, Yongqi Dong
Lane detection is crucial for vehicle localization which makes it the foundation for automated driving and many intelligent and advanced driving assistant systems. Available vision-based lane detection methods do not make full use of the valuable features and aggregate contextual information, especially the interrelationships between lane lines and other reg
Chen Xie, Blake Wilson, Zhenpeng Qin
Janus nanoparticles (JNPs) with heterogeneous compositions or interfacial properties can exhibit directional heating upon external excitation, such as laser radiation and magnetic field. This directional heating may be harnessed for new nanotechnology and biomedical applications. Understanding thermal transport and temperature control with JNP heating is cri
Noah Forman, Soumik Pal, Douglas Rizzolo, Matthias Winkel
Motivated by a down-up Markov chain on cladograms, David Aldous conjectured in 1999 that there exists a "diffusion on continuum trees" whose mass partitions at any finite number of branch points evolve as Wright-Fisher diffusions with some negative mutation rates, until some branch point disappears. Building on previous work on interval-partition-valued proc
Yucheng Li, Shun Wang, Chenghua Lin, Guerin Frank
One noticeable trend in metaphor detection is the embrace of linguistic theories such as the metaphor identification procedure (MIP) for model architecture design. While MIP clearly defines that the metaphoricity of a lexical unit is determined based on the contrast between its \textit{contextual meaning} and its \textit{basic meaning}, existing work does no
Md Mahfuz Ibn Alam, Sina Ahmadi, Antonios Anastasopoulos
Neural machine translation (NMT) systems exhibit limited robustness in handling source-side linguistic variations. Their performance tends to degrade when faced with even slight deviations in language usage, such as different domains or variations introduced by second-language speakers. It is intuitive to extend this observation to encompass dialectal variat
Vijeta Deshpande, Dan Pechi, Shree Thatte, Vladislav Lialin
In recent years, language models have drastically grown in size, and the abilities of these models have been shown to improve with scale. The majority of recent scaling laws studies focused on high-compute high-parameter count settings, leaving the question of when these abilities begin to emerge largely unanswered. In this paper, we investigate whether the
Colin Coane, Marco Romanelli, Giulia Dall'Osto, Rosa Di Felice
Electronic Energy Transfer (EET) between chromophores is fundamental in many natural light-harvesting complexes, serving as a critical step for solar energy funneling in photosynthetic plants and bacteria. The complicated role of the environment in mediating this process in natural architectures has been addressed by recent scanning tunneling microscope (STM
$SO(8)$ unification and the large-N theory of superconductor-insulator transition of two-dimensional Dirac fermions
cond-mat.str-elIgor F. Herbut, Subrata Mandal
Electrons on honeycomb or pi-flux lattices obey effective massless Dirac equation at low energies and at the neutrality point, and should suffer quantum phase transitions into various Mott insulators and superconductors at strong two-body interactions. We show that 35 out of 36 such order parameters that provide Lorentz-invariant mass-gaps to Dirac fermions
Geoffrey R. Harrison, Tobias Saule, R. Esteban Goetz, George N. Gibson
We demonstrate the generation of a train of attosecond XUV pulses that are in a superposition of wavefront states. Such superposition yields a high precision, self-referencing, common path XUV interferometer setup to produce pairs of spatially separated and independently controllable XUV pulses that are locked in phase and time with a temporal jitter of 3.5
Bhishma Dedhia, Michael Chang, Jake C. Snell, Thomas L. Griffiths
Large language models are few-shot learners that can solve diverse tasks from a handful of demonstrations. This implicit understanding of tasks suggests that the attention mechanisms over word tokens may play a role in analogical reasoning. In this work, we investigate whether analogical reasoning can enable in-context composition over composable elements of
Hussein Mozannar, Yuria Utsumi, Irene Y. Chen, Stephanie S. Gervasi
A high-risk pregnancy is a pregnancy complicated by factors that can adversely affect the outcomes of the mother or the infant. Health insurers use algorithms to identify members who would benefit from additional clinical support. This work presents the implementation of a real-world ML-based system to assist care managers in identifying pregnant patients at
Avinab Saha, Yu-Chih Chen, Chase Davis, Bo Qiu
We present the outcomes of a recent large-scale subjective study of Mobile Cloud Gaming Video Quality Assessment (MCG-VQA) on a diverse set of gaming videos. Rapid advancements in cloud services, faster video encoding technologies, and increased access to high-speed, low-latency wireless internet have all contributed to the exponential growth of the Mobile C
High-precision broadband linear polarimetry of early-type binaries IV. Binary system of DH Cephei in the open cluster of NGC 7380
astro-ph.SRYasir Abdul Qadir, Andrei V. Berdyugin, Vilppu Piirola, Takeshi Sakanoi
DH Cephei is a prominent O+O-type binary system in the NGC 7380 open cluster. Our high-precision multi-band polarimetry demonstrates synchronous linear polarization variations with the orbital period. Using Stokes parameters, we derived the binary system's orbital inclination, orientation, and rotation direction. To estimate interstellar polarization, we obs
Ho Chit Siu, Kevin Leahy, Makai Mann
Much of the recent work developing formal methods techniques to specify or learn the behavior of autonomous systems is predicated on a belief that formal specifications are interpretable and useful for humans when checking systems. Though frequently asserted, this assumption is rarely tested. We performed a human experiment (N = 62) with a mix of people who
Timothy Buttsworth, Artem Pulemotov
We prove local solvability of the Poisson equation with a positive or negative right-hand side for closed $G_2$-structures.
Ruixiang Tang, Dehan Kong, Longtao Huang, Hui Xue
Large language models (LLMs) have recently shown great potential for in-context learning, where LLMs learn a new task simply by conditioning on a few input-label pairs (prompts). Despite their potential, our understanding of the factors influencing end-task performance and the robustness of in-context learning remains limited. This paper aims to bridge this
Michele Lohr, Laurent Younes
A multivariate regression model of affine and diffeomorphic transformation sequences - FineMorphs - is presented. Leveraging concepts from shape analysis, model states are optimally "reshaped" by diffeomorphisms generated by smooth vector fields during learning. Affine transformations and vector fields are optimized within an optimal control setting, and the
Jean-Louis Krivine
A remark on the proof that the Grothendieck constant satisfies $K_G < \pi/(2\ln(1+\sqrt{2}))$.
Wonoo Choo, Erkan Kayacan
This paper develops computationally efficient data-driven model predictive control (MPC) for Agile quadrotor flight. Agile quadrotors in high-speed flights can experience high levels of aerodynamic effects. Modeling these turbulent aerodynamic effects is a cumbersome task and the resulting model may be overly complex and computationally infeasible. Combining
Reliability Evaluation of Phasor Measurement Unit Considering Failure of Hardware and Software Using Fuzzy Approach
eess.SPEvan Carollo, Zikai Xu
The wide-area measurement system (WAMS) consists of the future power system, increasing geographical sprawl which is linked by the Phasor measurement unit(PMU). Thus, the failure of PMU will cause severe results, such as a blackout of the power system. In this paper, the reliability model of PMU is considered both hardware and software, where it gives a char
Vaibhav Saxena, Kamal Rahimi Malekshan, Linh Tran, Yotto Koga
6-DoF pose estimation is an essential component of robotic manipulation pipelines. However, it usually suffers from a lack of generalization to new instances and object types. Most widely used methods learn to infer the object pose in a discriminative setup where the model filters useful information to infer the exact pose of the object. While such methods o
Sonny Achten, Arun Pandey, Hannes De Meulemeester, Bart De Moor
We propose a unifying setting that combines existing restricted kernel machine methods into a single primal-dual multi-view framework for kernel principal component analysis in both supervised and unsupervised settings. We derive the primal and dual representations of the framework and relate different training and inference algorithms from a theoretical per
Boyuan Chen, Chuning Zhu, Pulkit Agrawal, Kaiqing Zhang
Model-free reinforcement learning algorithms have exhibited great potential in solving single-task sequential decision-making problems with high-dimensional observations and long horizons, but are known to be hard to generalize across tasks. Model-based RL, on the other hand, learns task-agnostic models of the world that naturally enables transfer across dif
A Reissner-Mindlin plate formulation using symmetric Hu-Zhang elements via polytopal transformations
math.NAAdam Sky, Michael Neunteufel, Jack S. Hale, Andreas Zilian
In this work we develop new finite element discretisations of the shear-deformable Reissner--Mindlin plate problem based on the Hellinger-Reissner principle of symmetric stresses. Specifically, we use conforming Hu-Zhang elements to discretise the bending moments in the space of symmetric square integrable fields with a square integrable divergence $\boldsym
Fatemeh Rezaei, Diluka Galappaththige, Chintha Tellambura, Amine Maaref
Current backscatter channel estimators employ an inefficient silent pilot transmission protocol, where tags alternate between silent and active states. To enhance performance, we propose a novel approach where tags remain active simultaneously throughout the entire training phase. This enables a one-shot estimation of both the direct and cascaded channels an
The Araucaria Project, :, G. Pietrzyński, W. Gieren
The book consists of a number of short articles that present achievements of the Araucaria members, collaborators, and friends, in various aspects of distance determinations and related topics. It celebrates the 20-year anniversary of the Araucaria Project, acknowledges the people who worked for its success, and popularises our methods and results among broa
NASimEmu: Network Attack Simulator & Emulator for Training Agents Generalizing to Novel Scenarios
cs.CRJaromír Janisch, Tomáš Pevný, Viliam Lisý
Current frameworks for training offensive penetration testing agents with deep reinforcement learning struggle to produce agents that perform well in real-world scenarios, due to the reality gap in simulation-based frameworks and the lack of scalability in emulation-based frameworks. Additionally, existing frameworks often use an unrealistic metric that meas
Hamoon Jafarian, Faisal Z. Qureshi
Human pose and shape estimation methods continue to suffer in situations where one or more parts of the body are occluded. More importantly, these methods cannot express when their predicted pose is incorrect. This has serious consequences when these methods are used in human-robot interaction scenarios, where we need methods that can evaluate their predicti
Ketaki Joshi, Raghavendra Pradyumna Pothukuchi, Andre Wibisono, Abhishek Bhattacharjee
Continual learning on sequential data is critical for many machine learning (ML) deployments. Unfortunately, LSTM networks, which are commonly used to learn on sequential data, suffer from catastrophic forgetting and are limited in their ability to learn multiple tasks continually. We discover that catastrophic forgetting in LSTM networks can be overcome in
T. de Jaeger, L. Galbany
The use of multiple independent methods with their own systematic uncertainties is crucial for resolving the ongoing tension between local and distant measurements of the Hubble constant ($H_{0}$). While type Ia supernovae (SNe Ia) have historically been the most widely used distance indicators, recent studies have shown that type II supernovae (SNe II) can
Validating phase-space methods with tensor networks in two-dimensional spin models with power-law interactions
quant-phSean R. Muleady, Mingru Yang, Steven R. White, Ana Maria Rey
Using a recently developed extension of the time-dependent variational principle for matrix product states, we evaluate the dynamics of 2D power-law interacting XXZ models, implementable in a variety of state-of-the-art experimental platforms. We compute the spin squeezing as a measure of correlations in the system, and compare to semiclassical phase-space c
Ahmad Al-Omari, Murad Ozcog, Santanu Acharjee
The main purpose of this paper is to introduce and study the primal-proximity spaces. Also, we define two new operators via primal proximity spaces and investigate some of their fundamental properties. In addition, we obtain a new topology, which is weaker than old one, via these new operators. Moreover, we not only discuss some of their properties but also
Jameson Cahill, Joseph W. Iverson, Dustin G. Mixon
Consider the quotient of a Hilbert space by a subgroup of its automorphisms. We study whether this orbit space can be embedded into a Hilbert space by a bilipschitz map, and we identify constraints on such embeddings.
Zehui Lu, Shaoshuai Mou
Generalized from the concept of consensus, this paper considers a group of edge agreements, i.e. constraints defined for neighboring agents, in which each pair of neighboring agents is required to satisfy one edge agreement constraint. Edge agreements are defined locally to allow more flexibility than a global consensus. This work formulates a multi-agent op
Sarah Cannon
In the United States, regions are frequently divided into districts for the purpose of electing representatives. How the districts are drawn can affect who's elected, and drawing districts to give an advantage to a certain group is known as gerrymandering. It can be surprisingly difficult to detect gerrymandering, but one algorithmic method is to compare a c
R. Cerroni, S. Dell'Oro, A. Formicola, S. Ghislandi
$^{180m}$Ta is the longest-lived metastable state presently known. Its decay has not been observed yet. In this work, we report a new result on the decay of \mTa obtained with a $2015.12$-g tantalum sample measured for $527.7$ d with an ultra-low background HPGe detector in the STELLA laboratory of the Laboratori Nazionali del Gran Sasso (LNGS), in Italy. Be
Asymptotically locally flat and AdS higher-dimensional black holes of Einstein-Horndeski-Maxwell gravity in the light of EHT observations: shadow behavior and deflection angle
gr-qcKourosh Nozari, Sara Saghafi
Unification of gravity with other interactions, achieving the ultimate framework of quantum gravity, and fundamental problems in particle physics and cosmology motivate to consider extra spatial dimensions. The impact of these extra dimensions on the modified theories of gravity has attracted a lot of attention. One way to examine how extra dimensions affect
Maxime Ramzi
We study the notion of \emph{separable algebras} in the context of symmetric monoidal stable $\infty$-categories. In the first part of this paper, we compare this context to that of tensor-triangulated categories and show that separable algebras and their modules in a symmetric monoidal stable $\infty$-category are, in large parts, controlled by the (tensor-
Jinqi Xiao, Miao Yin, Yu Gong, Xiao Zang
Attention-based vision models, such as Vision Transformer (ViT) and its variants, have shown promising performance in various computer vision tasks. However, these emerging architectures suffer from large model sizes and high computational costs, calling for efficient model compression solutions. To date, pruning ViTs has been well studied, while other compr