Skip to content

December 2024 arXiv papers — page 100

Showing 9,90110,000 of 20,868 papers

  1. Xingchi Chen, Zhuoran Zheng, Xuerui Li, Yuying Chen

    With the continuous improvement of device imaging resolution, the popularity of Ultra-High-Definition (UHD) images is increasing. Unfortunately, existing methods for fusing multi-exposure images in dynamic scenes are designed for low-resolution images, which makes them inefficient for generating high-quality UHD images on a resource-constrained device. To al

  2. Benjamin Doerr, Martin S. Krejca, Günter Rudolph

    Randomized search heuristics have been applied successfully to a plethora of problems. This success is complemented by a large body of theoretical results. Unfortunately, the vast majority of these results regard problems with binary or continuous decision variables -- the theoretical analysis of randomized search heuristics for unbounded integer domains is

  3. Dexter Le, Aybars Yunusoglu, Karn Tiwari, Murat Isik

    In the evolving landscape of transportation systems, integrating Large Language Models (LLMs) offers a promising frontier for advancing intelligent decision-making across various applications. This paper introduces a novel 3-dimensional framework that encapsulates the intersection of applications, machine learning methodologies, and hardware devices, particu

  4. Chengyue Wang, Haicheng Liao, Bonan Wang, Yanchen Guan

    Accurate trajectory prediction is essential for the safety and efficiency of autonomous driving. Traditional models often struggle with real-time processing, capturing non-linearity and uncertainty in traffic environments, efficiency in dense traffic, and modeling temporal dynamics of interactions. We introduce NEST (Neuromodulated Small-world Hypergraph Tra

  5. Abdelbaki Souid, Mohamed Hamroun, Soufiene Ben Othman, Hedi Sakli

    Pulmonary pathologies are a significant global health concern, often leading to fatal outcomes if not diagnosed and treated promptly. Chest radiography serves as a primary diagnostic tool, but the availability of experienced radiologists remains limited. Advances in Artificial Intelligence (AI) and machine learning, particularly in computer vision, offer pro

  6. Zheng Fang, Ke Ye, Yaofang Liu, Gongzhe Li

    Point clouds or depth images captured by current RGB-D cameras often suffer from low resolution, rendering them insufficient for applications such as 3D reconstruction and robots. Existing point cloud super-resolution (PCSR) methods are either constrained by geometric artifacts or lack attention to edge details. To address these issues, we propose an edge-gu

  7. Daiki Shirafuji, Makoto Takenaka, Shinya Taguchi

    The use of language models (LMs) has increased considerably in recent years, and the biases and stereotypes in training data that are reflected in the LM outputs are causing social problems. In this paper, inspired by the task arithmetic, we propose the ``Bias Vector'' method for the mitigation of these LM biases. The Bias Vector method does not require manu

  8. Shuai Zhou, Shizhe Zhao, Zhongqiang Ren

    Multi-Agent Path Finding (MAPF) seeks collision-free paths for multiple agents from their respective starting locations to their respective goal locations while minimizing path costs. Although many MAPF algorithms were developed and can handle up to thousands of agents, they usually rely on the assumption that each action of the agent takes a time unit, and

  9. Beibei Li, Yutian Chi, Yuming Wang

    This study introduces a novel approach that integrates the magnetic field data correction from the Tianwen-1 Mars mission with a neural network architecture constrained by physical principles derived from Maxwell's equation equations. By employing a Transformer based model capable of efficiently handling sequential data, the method corrects measurement anoma

  10. Jangho Kim, Thomas Luu, Wolfgang Unger

    In our previous studies [1, 2], we confirmed that a quantum annealer can be used for importance sampling of gauge theories. In this paper, we extend the previous results to larger 2-dimensional and 4-dimensional lattices to generate ensembles for U(3) gauge theory in the strong coupling limit. We make use of the D-Wave quantum annealer to generate histograms

  11. Thierry Dana-Picard

    Hyperbolism of a given curve with respect to a point and a line is an interesting construct, a special kind of geometric locus, not frequent in the literature. While networking between two different kinds of mathematical software, we explore various cases, involving quartics, among them the so-called Kuelp quartic and topologically equivalent curves, and als

  12. Zhengyu Yin

    In this paper, inspired by the elegant work of Good and Meddaugh \cite{GM} and the graph models for zero-dimensional systems developed by several authors, like Gambaudo and Martens \cite{GM06}, Shimomura \cite{Sh14}. We try to discover a connection among some objects, such as finite directed graph, shift of finite type and shadowing property by employing the

  13. Hangyu Zhu, Yuxiang Fan, Zhenping Xie

    Federated learning (FL) is a privacy preserving machine learning paradigm designed to collaboratively learn a global model without data leakage. Specifically, in a typical FL system, the central server solely functions as an coordinator to iteratively aggregate the collected local models trained by each client, potentially introducing single-point transmissi

  14. Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis

    Predicting future dynamics is crucial for applications like autonomous driving and robotics, where understanding the environment is key. Existing pixel-level methods are computationally expensive and often focus on irrelevant details. To address these challenges, we introduce DINO-Foresight, a novel framework that operates in the semantic feature space of pr

  15. Lillian Wassim, Kamal Mohamed, Ali Hamdi

    We propose LLM-DaaS, a novel Drone-as-a-Service (DaaS) framework that leverages Large Language Models (LLMs) to transform free-text user requests into structured, actionable DaaS operation tasks. Our approach addresses the key challenge of interpreting and structuring natural language input to automate drone service operations under uncertain conditions. The

  16. Jangyeong Jeon, Sangyeon Cho, Dongjoon Lee, Changhee Lee

    Pediatric Emergency Department (PED) overcrowding presents a significant global challenge, prompting the need for efficient solutions. This paper introduces the BioBridge framework, a novel approach that applies Natural Language Processing (NLP) to Electronic Medical Records (EMRs) in written free-text form to enhance decision-making in PED. In non-English s

  17. Jonatan Piasetzky, Yehonatan Drori, Yuval Warshavski, Amit Rotem

    Directional couplers are essential components in integrated photonics. Given their widespread use, accurate characterization of directional couplers is crucial for ensuring optimal performance. However, it is challenging due to the coupling between fibers and waveguides, which is highly sensitive to alignment and fabrication imperfections. To address these c

  18. Matthias Lanzinger, Cem Okulmus, Reinhard Pichler, Alexander Selzer

    Hypertree decompositions provide a way to evaluate Conjunctive Queries (CQs) in polynomial time, where the exponent of this polynomial is determined by the width of the decomposition. In theory, the goal of efficient CQ evaluation therefore has to be a minimisation of the width. However, in practical settings, it turns out that there are also other propertie

  19. Peirong Zhang, Lianwen Jin

    Currently, the prevalence of online handwriting has spurred a critical need for effective retrieval systems to accurately search relevant handwriting instances from specific writers, known as online writer retrieval. Despite the growing demand, this field suffers from a scarcity of well-established methodologies and public large-scale datasets. This paper ta

  20. Alessio Di Santo, Walter Tiberti, Dajana Cassioli

    Quantum secret sharing (QSS) is a cryptographic protocol that leverages quantum mechanics to distribute a secret among multiple parties. With respect to the classical counterpart, in QSS the secret is encoded into quantum states and shared by a dealer such that only an authorized subsets of participants, i.e., the players, can reconstruct it. Several state-o

  21. Luca M. Hartmann, Orcun Karaca, Tinus Dorfling, Tobias Geyer

    This paper formulates a semidefinite programming relaxation for a long horizon direct-torque finite-control-set model predictive control problem. In parallel with this relaxation, a conventional branch-and-bound algorithm tailored for the original problem, but with an iteration limit to restrict its computational burden, is also solved. An input sequence can

  22. Sanjay Chakraborty, Ibrahim Delibasoglu, Fredrik Heintz

    Time series forecasting is a crucial challenge with significant applications in areas such as weather prediction, stock market analysis, and scientific simulations. This paper introduces an embedded decomposed transformer, 'EDformer', for multivariate time series forecasting tasks. Without altering the fundamental elements, we reuse the Transformer architect

  23. Linhui Chen, Qi Liu, Xiewei Tan, Yuxin Wang

    For any $\lambda>1, R_\lambda^2$ is Bana\'s-Fr\k{a}czek space, the exact value of the skew generalized von Neumann-Jordan constant $C_{\mathrm{NJ}}^p\left(\xi, \eta, R_\lambda^2\right)$ is calculated. By careful calculations, $C_{\mathrm{NJ}}^p\left(\xi, \eta, R_\lambda^2\right)=\frac{(\xi+\eta)^p+\left[(\eta+\xi)^2-\frac{4 \xi \eta}{\lambda^2}\right]^{p / 2

  24. Yu Kang, Xianghui Sun, Liangyu Chen, Wei Zou

    Generating Chain-of-Thought (CoT) before deriving the answer can effectively improve the reasoning capabilities of large language models (LLMs) and significantly improve the accuracy of the generated answer. However, in most cases, the length of the generated CoT is much longer than the desired final answer, which results in additional decoding costs. Furthe

  25. Maria Tzelepi, Vasileios Mezaris

    In this paper we deal with image classification tasks using the powerful CLIP vision-language model. Our goal is to advance the classification performance using the CLIP's image encoder, by proposing a novel Large Multimodal Model (LMM) based regularization method. The proposed method uses an LMM to extract semantic descriptions for the images of the dataset

  26. Patrik Demjan, N. C. Snaith

    We consider the $n$-correlation of eigenvalues of random unitary matrices in the alternative form that is not the tidy determinant common in random matrix theory, but rather the expression derived from averages of ratios of characteristic polynomials in a method that can be mimicked in number theoretical calculations of the correlations of zeros of $L$-funct

  27. Phokion G. Kolaitis, Nina Pardal, Jonni Virtema, Jef Wijsen

    We embark on a study of the consistent answers of queries over databases annotated with values from a naturally ordered positive semiring. In this setting, the consistent answers of a query are defined as the minimum of the semiring values that the query takes over all repairs of an inconsistent database. The main focus is on self-join free conjunctive queri

  28. Dipanwita Thakur, Antonella Guzzo, Giancarlo Fortino, Sajal K. Das

    This paper proposes a novel federated algorithm that leverages momentum-based variance reduction with adaptive learning to address non-convex settings across heterogeneous data. We intend to minimize communication and computation overhead, thereby fostering a sustainable federated learning system. We aim to overcome challenges related to gradient variance, w

  29. Gabriele Bressanini, Farhan Hanif, Hyukjoon Kwon, M. S. Kim

    We introduce the concept of quantum observables over time (QOOT), an operator that jointly describes two observables at two distinct time points, as a dual of the quantum state over time formalism. We provide a full characterization of the conditions under which a QOOT can be properly defined, via a no-go theorem. We use QOOTs to establish a notion of time-r

  30. Tianyi Yin, Jingwei Wang, Yunlong Ma, Han Wang

    Encoding time series into tokens and using language models for processing has been shown to substantially augment the models' ability to generalize to unseen tasks. However, existing language models for time series forecasting encounter several obstacles, including aliasing distortion and prolonged inference times, primarily due to the limitations of quantiz

  31. Gaurav Aggarwal, Anish Ghosh

    We provide the first known upper bounds for the packing dimension of weighted singular and weighted $\omega$-singular matrices. We also prove upper bounds for these sets when intersected with fractal subsets. The latter results, even in the unweighted setting, are already new for matrices. Further, even for row vectors, our results enlarge the class of fract

  32. Nikhil Kapila, Julian Glattki, Tejas Rathi

    Convolutional Neural Networks (CNNs) have been the standard for image classification tasks for a long time, but more recently attention-based mechanisms have gained traction. This project aims to compare traditional CNNs with attention-augmented CNNs across an image classification task. By evaluating and comparing their performance, accuracy and computationa

  33. Ryuhaerang Choi, Taehan Kim, Subin Park, Jennifer G Kim

    Eating disorders (ED) are complex mental health conditions that require long-term management and support. Recent advancements in large language model (LLM)-based chatbots offer the potential to assist individuals in receiving immediate support. Yet, concerns remain about their reliability and safety in sensitive contexts such as ED. We explore the opportunit

  34. Francesco Malaspina

    We introduce the notion of primitive Ulrich bundle in a smooth projective variety. We motivate this notion and give a cohomological characterization in the case of the degree $6$ flag threefold and rational normal scrolls. Finally we propose a few open problems.

  35. V. A. Berezin, I. D. Ivanova, A. E. Kuprina

    The phenomenological description of the cosmological particle production in the framework of the induced gravity is investigated. It appears that the source terms with the particle number density in the creation law can be interpreted as the invisible part of the Universe. It is shown that there is a gauge that restores the General Relativity in which our mo

  36. Wei Chen, Guo Ye, Yakun Wang, Zhao Zhang

    Unsupervised Graph Domain Adaptation (UGDA) seeks to bridge distribution shifts between domains by transferring knowledge from labeled source graphs to given unlabeled target graphs. Existing UGDA methods primarily focus on aligning features in the latent space learned by graph neural networks (GNNs) across domains, often overlooking structural shifts, resul

  37. Amelie Wührl, Roman Klinger

    In fact-checking, structure and phrasing of claims critically influence a model's ability to predict verdicts accurately. Social media content in particular rarely serves as optimal input for verification systems, which necessitates pre-processing to extract the claim from noisy context before fact checking. Prior work suggests extracting a claim representat

  38. Tao Meng, Wei Ai, Jianbin Li, Ze Wang

    Text representation learning is significant as the cornerstone of natural language processing. In recent years, graph contrastive learning (GCL) has been widely used in text representation learning due to its ability to represent and capture complex text information in a self-supervised setting. However, current mainstream graph contrastive learning methods

  39. Dihong Huang

    Sequential inspection is a technique employed to monitor product quality during the production process. For smaller batch sizes, the Acceptable Quality Limit(AQL) inspection theory is typically applied, whereas for larger batch sizes, the Poisson distribution is commonly utilized to determine the sample size and rejection thresholds. However, due to the fact

  40. Kaixuan Wang, Lin Qi, Shiyu Qin, Kai Luo

    Photometric stereo (PS) endeavors to ascertain surface normals using shading clues from photometric images under various illuminations. Recent deep learning-based PS methods often overlook the complexity of object surfaces. These neural network models, which exclusively rely on photometric images for training, often produce blurred results in high-frequency

  41. Miguel Correia, Holmfridur S. Hannesdottir, Giulia Isabella, Anna M. Wolz

    These lecture notes explain how classical gravitational physics emerges from scattering amplitudes. We emphasize the role of different kinematic regimes in probing various aspects of bound and unbound problems, as illustrated by the Hydrogen atom example. Classical predictions of General Relativity, such as the Shapiro time delay and perihelion precession, e

  42. Arnaud Le Fèvre, Abdelouahad Chbihi, Quentin Fable, Tom Génard

    A new method, based on comparing isotopic yield ratios measured at forward and sideward polar angles and on cross-bombarding heavy nuclei with different neutron-to-proton ratios, is used to quantify the stopping power of nuclear matter in heavy-ion collisions. For central collisions of isotopically separated $^{124,129}$Xe+$^{112,124}$Sn at 100~MeV/nucleon b

  43. Tianyi Chen, Atsushi Miyauchi, Charalampos E. Tsourakakis

    Given a network $G=(V,E)$, where each node $v$ is associated with a vector $\boldsymbol{p}_v \in \mathbb{R}^d$ representing its opinion about $d$ different topics, how can we uncover subsets of nodes that not only exhibit exceptionally high density but also possess positively aligned opinions on multiple topics? In this paper we focus on this novel algorithm

  44. Nour Jamoussi, Giuseppe Serra, Photios A. Stavrou, Marios Kountouris

    Federated learning (FL) is a widely used and impactful distributed optimization framework that achieves consensus through averaging locally trained models. While effective, this approach may not align well with Bayesian inference, where the model space has the structure of a distribution space. Taking an information-geometric perspective, we reinterpret FL a

  45. LHCb collaboration, R. Aaij, A. S. W. Abdelmotteleb, C. Abellan Beteta

    The first test of lepton flavor universality between muons and electrons using $B^+ \to K^+\pi^+\pi^-\ell^+\ell^-$ ($\ell=e,\mu$) decays is presented. The measurement is performed with data from proton-proton collisions collected by the LHCb experiment at center-of-mass energies of 7, 8, and 13 TeV, corresponding to an integrated luminosity of $9\mathrm{fb}^

  46. ATLAS Collaboration

    A set of measurements for the production of a $W$-boson in association with high-transverse-momentum jets is presented using 140 fb$^{-1}$ of proton-proton collision data at a centre-of-mass energy of $\sqrt{s}=13$ TeV collected by the ATLAS detector at the LHC. The measurements are performed in final states in which the $W$-boson decays into an electron or

  47. Anasse Boutayeb, Iyad Lahsen-cherif, Ahmed El Khadimi

    In recent years, Geospatial Artificial Intelligence (GeoAI) has gained traction in the most relevant research works and industrial applications, while also becoming involved in various fields of use. This paper offers a comprehensive review of GeoAI as a synergistic concept applying Artificial Intelligence (AI) methods and models to geospatial data. A prelim

  48. Waleed El-Geresy, Christos Papavassiliou, Deniz Gündüz

    In this paper, we examine the problem of information storage on memristors affected by resistive drift noise under energy constraints. We introduce a novel, fundamental trade-off between the information lifetime of memristive states and the energy that must be expended to bring the device into a particular state. We then treat the storage problem as one of c

  49. Marco Aiello, Ilche Georgievski

    These are notes for lectures presented at the University of Stuttgart that provide an introduction to key concepts and techniques in AI Planning. Artificial Intelligence Planning, also known as Automated Planning, emerged somewhere in 1966 from the need to give autonomy to a wheeled robot. Since then, it has evolved into a flourishing research and developmen

  50. Sandro Donato, Simone Caputo, Luca Brombal, Bruno Golosio

    Three different computed tomography (CT) reconstruction algorithms: Filtered Back Projection (FBP), Unified Tomographic Reconstruction (UTR) and customized Simultaneous Algebraic Reconstruction Technique (cSART), have been systematically compared and evaluated using experimental data from CT scans of ten fresh mastectomy samples collected at the Imaging and

  51. Guoyu Hu, Yuncheng Wu, Gang Chen, Tien Tuan Anh Dinh

    Model inference systems are essential for implementing end-to-end data analytics pipelines that deliver the benefits of machine learning models to users. Existing cloud-based model inference systems are costly, not easy to scale, and must be trusted in handling the models and user request data. Serverless computing presents a new opportunity, as it provides

  52. Wei Zhang, Weiquan Yan, Yun Zhao, Wenxiang Cheng

    Neuromorphic vision sensors, such as the dynamic vision sensor (DVS) and spike camera, have gained increasing attention in recent years. The spike camera can detect fine textures by mimicking the fovea in the human visual system, and output a high-frequency spike stream. Real-time high-quality vision reconstruction from the spike stream can build a bridge to

  53. Yiren Song, Pei Yang, Hai Ci, Mike Zheng Shou

    Recently, zero-shot methods like InstantID have revolutionized identity-preserving generation. Unlike multi-image finetuning approaches such as DreamBooth, these zero-shot methods leverage powerful facial encoders to extract identity information from a single portrait photo, enabling efficient identity-preserving generation through a single inference pass. H

  54. Frances Yung, Vera Demberg

    Interpreting implicit discourse relations involves complex reasoning, requiring the integration of semantic cues with background knowledge, as overt connectives like because or then are absent. These relations often allow multiple interpretations, best represented as distributions. In this study, we compare two established methods that crowdsource English im

  55. Janmejoy Sarkar, Rushikesh Deogaonkar, Ravi Kesharwani, Sreejith Padinhatteeri

    The Solar Ultraviolet Imaging Telescope (SUIT) on board the Aditya-L1 mission is designed to observe the Sun across 200-400 nm wavelength. The telescope used 16 dichroic filters tuned at specific wavelengths in various combinations to achieve its science goals. For accurate measurements and interpretation, it is important to characterize these filters for sp

  56. Olga Kulitckaya, Alexander Gorfer, Elena Petrishcheva, Bengü Tas

    Tracer diffusion of Na in natural alkali feldspars including sanidine, adularia and orthoclase with different Na:K ratios is measured using the radiotracer technique and applying the 22Na radioisotope. The tracer diffusion measurements along the crystallographic directions perpendicular to (001) and (010) in alularia feldspar reveled a slight (within a facto

  57. Zhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang

    Historical documents encompass a wealth of cultural treasures but suffer from severe damages including character missing, paper damage, and ink erosion over time. However, existing document processing methods primarily focus on binarization, enhancement, etc., neglecting the repair of these damages. To this end, we present a new task, termed Historical Docum

  58. Alexandre C. Orthey, Alexander Streltsov

    Quantum realism, as introduced by Bilobran and Angelo [EPL 112, 40005 (2015)], states that projective measurements in quantum systems establish the reality of physical properties, even in the absence of a revealed outcome. This framework provides a nuanced perspective on the distinction between classical and quantum notions of realism, emphasizing the contex

  59. Juncheng Zou

    Accurate human motion prediction is crucial for safe human-robot collaboration but remains challenging due to the complexity of modeling intricate and variable human movements. This paper presents Parallel Multi-scale Incremental Prediction (PMS), a novel framework that explicitly models incremental motion across multiple spatio-temporal scales to capture su

  60. Yuyang Tao, Shufei Ge

    The Mapper algorithm is an essential tool for visualizing complex, high dimensional data in topology data analysis (TDA) and has been widely used in biomedical research. It outputs a combinatorial graph whose structure implies the shape of the data. However,the need for manual parameter tuning and fixed intervals, along with fixed overlapping ratios may impe

  61. Filippo Marini, Margherita Porcelli, Elisa Riccietti

    In this paper, we propose a multilevel stochastic framework for the solution of nonconvex unconstrained optimization problems. The proposed approach uses random regularized first-order models that exploit an available hierarchical description of the problem, being either in the classical variable space or in the function space, meaning that different levels

  62. Changhai Zhou, Yuhua Zhou, Shijie Han, Qian Qiao

    The rise of large language models (LLMs) has significantly advanced various natural language processing (NLP) tasks. However, the resource demands of these models pose substantial challenges. Structured pruning is an effective approach to reducing model size, but it often results in significant accuracy degradation, necessitating parameter updates to adapt.

  63. Fan Xu, Yutong Yu

    We study quantum cluster algebras from marked surfaces without punctures. We express the quantum cluster variables in terms of the canonical submodules. As a byproduct, we obtain the positivity for this class of quantum cluster algebra.

  64. Elsa Ducrot, Pierre-Olivier Lagage, Michiel Min, Michael Gillon

    The first JWST/MIRI photometric observations of TRAPPIST-1 b allowed for the detection of the thermal emission of the planet at 15 $\mu m$, suggesting that the planet could be a bare rock with a zero albedo and no redistribution of heat. These observations at 15 $\mu m$ were acquired as part of GTO time that included a twin program at 12.8 $\mu m$ in order t

  65. Patrik L. Ferrari, Min Liu

    Backwards geodesics for TASEP were introduced in [Fer18]. We consider flat initial conditions and show that under proper scaling its end-point converges to maximizer argument of the Airy$_2$ process minus a parabola. We generalize its definition to generic non-integrable models including ASEP and speed changed ASEP (call it quasi-geodesics). We numerically v

  66. Diana Bar-Or Nirman, Ariel Weizman, Amos Azaria

    While Large Language Models (LLMs) have become central tools in various fields, they often provide inaccurate or false information. This study examines user preferences regarding falsehood responses from LLMs. Specifically, we evaluate preferences for LLM responses where false statements are explicitly marked versus unmarked responses and preferences for con

  67. Irek Mukhamedshin, Pawel Wzietek, Fabrice Bert, Philippe Mendels

    A magnetic Weyl semimetal presents the intriguing possibility of controlling topological properties through magnetic order. The kagome compound \CoSnS~has emerged as one of the most thoroughly characterized magnetic Weyl semimetals, yet the potential coexistence of a ferromagnetic state below $T_c$ = 172~K with a non-collinear antiferromagnetic phase or a gl

  68. Yalun Zhang, Longwen Zhou

    Non-Hermitian phenomena, such as exceptional points, non-Hermitian skin effects, and topologically nontrivial phases have attracted continued attention. In this work, we reveal how interactions and nonreciprocal hopping could collectively influence the behavior of two interacting bosons on quasiperiodic lattices. Focusing on the Bose-Hubbard model with Aubry

  69. Lukas Mauth

    We will prove an infinite family of asymptotic formulas for the logarithm of certain two-colored partitions. An infinite sub-family of these asymptotics was posed as a conjecture by Guadalupe.

  70. Muhammet Furkan Ilaslan, Ali Koksal, Kevin Qinhong Lin, Burak Satar

    Large Language Model (LLM)-based agents have shown promise in procedural tasks, but the potential of multimodal instructions augmented by texts and videos to assist users remains under-explored. To address this gap, we propose the Visually Grounded Text-Video Prompting (VG-TVP) method which is a novel LLM-empowered Multimodal Procedural Planning (MPP) framew

  71. Wenxiao Fan, Kan Li

    Noisy labels can negatively impact the performance of deep neural networks. One common solution is label refurbishment, which involves reconstructing noisy labels through predictions and distributions. However, these methods may introduce problematic semantic associations, a phenomenon that we identify as Semantic Contamination. Through an analysis of Robust

  72. Ixandra Achitouv, Vincent Lahoche, Dine Ousmane Samary, Parham Radpay

    In this paper, we consider a renormalization group perspective on the quantum dynamics of a particle moving in the Euclidean $\mathbb{R}^N$ space through the complex landscape provided by a disordered Hamiltonian of type $2+p$. We focus on the large $N$ limit, where the coarse-graining procedure is unconventional: it is based on the Wigner spectrum of the ra

  73. Pan Wang, Qiang Zhou, Yawen Wu, Tianlong Chen

    Multimodal Sentiment Analysis (MSA) leverages heterogeneous modalities, such as language, vision, and audio, to enhance the understanding of human sentiment. While existing models often focus on extracting shared information across modalities or directly fusing heterogeneous modalities, such approaches can introduce redundancy and conflicts due to equal trea

  74. Nuowei Liu, Changzhi Sun, Tao Ji, Junfeng Tian

    Current Large Language Models (LLMs) for understanding proteins primarily treats amino acid sequences as a text modality. Meanwhile, Protein Language Models (PLMs), such as ESM-2, have learned massive sequential evolutionary knowledge from the universe of natural protein sequences. Furthermore, structure-based encoders like ProteinMPNN learn the structural i

  75. Yasir Mahmood, Markus Hecher, Axel-Cyrille Ngonga Ngomo

    The connection between inconsistent databases and Dung's abstract argumentation framework has recently drawn growing interest. Specifically, an inconsistent database, involving certain types of integrity constraints such as functional and inclusion dependencies, can be viewed as an argumentation framework in Dung's setting. Nevertheless, no prior work has ex

  76. Chunyan Lu, Liangliang Ren, Jiamao Lin, Wenjun Huang

    Spider pulsars constitute a distinct subset within the domain of radio millisecond pulsars, divided further into the categories of black widows and redbacks. Evident across multiple wavelengths, these pulsars manifest periodic variations and reside within binary systems. Investigating and discovering additional spider-type pulsars carries significant implica

  77. Javier García Gilabert, Carlos Escolano, Audrey Mash, Xixian Liao

    We introduce MT-LENS, a framework designed to evaluate Machine Translation (MT) systems across a variety of tasks, including translation quality, gender bias detection, added toxicity, and robustness to misspellings. While several toolkits have become very popular for benchmarking the capabilities of Large Language Models (LLMs), existing evaluation tools of

  78. Ruiyang Xia, Guanjun Gao, Zanshan Zhao, Haoyu Wang

    The enhanced Gaussian noise (EGN) model, which accounts for inter-channel stimulated Raman scattering (ISRS), has been extensively utilized for evaluating nonlinear interference (NLI) within the C+L band. Compared to closed-form expressions and machine learning-based NLI evaluation models, it demonstrates broader applicability and its accuracy is not depende

  79. Louis Gabarra

    The Euclid telescope, launched from Cape Canaveral on July 1st, 2023, is dedicated to studying dark matter and dark energy from its orbit at the Sun-Earth Lagrangian point L2. It is equipped with two instruments: the visual imager (VIS) and the Near-Infrared Spectrometer and Photometer (NISP). The Euclid Wide Survey (Scaramella et al. 2022) will cover approx

  80. Ferdinand V. Stoye, Annika Hoyer, Roland Langrock

    New types of high-resolution animal movement data allow for increasingly comprehensive biological inference, but method development to meet the statistical challenges associated with such data is lagging behind. In this contribution, we extend the commonly applied hidden Markov models for step lengths and turning angles to address the specific requirements p

  81. David Feldstein-Bofill, Zhenhai Sun, Casper Wied, Shikhar Singh

    The development of quantum circuits based on hybrid superconductor-semiconductor Josephson junctions holds promise for exploring their mesoscopic physics and for building novel superconducting devices. The gate-tunable superconducting transmon qubit (gatemon) is the paradigmatic example of such a superconducting circuit. However, gatemons typically suffer fr

  82. Tatsuya Shibata, Michael Conrad Koch, Iason Papaioannou, Kazunori Fujisawa

    Detection of abrupt spatial changes in physical properties representing unique geometric features such as buried objects, cavities, and fractures is an important problem in geophysics and many engineering disciplines. In this context, simultaneous spatial field and geometry estimation methods that explicitly parameterize the background spatial field and the

  83. Bingwen Hu, Heng Liu, Zhedong Zheng, Ping Liu

    Convolutional Neural Networks (CNNs) have significantly advanced Image Super-Resolution (SR), yet most CNN-based methods rely solely on pixel-based transformations, often leading to artifacts and blurring, particularly under severe downsampling rates (\eg, 8$\times$ or 16$\times$). The recently developed text-guided SR approaches leverage textual description

  84. Svetlana Pavlitska, Enrico Eisen, J. Marius Zöllner

    Vulnerability to adversarial attacks is a well-known deficiency of deep neural networks. Larger networks are generally more robust, and ensembling is one method to increase adversarial robustness: each model's weaknesses are compensated by the strengths of others. While an ensemble uses a deterministic rule to combine model outputs, a mixture of experts (MoE

  85. Mohammed Srati

    In this paper, we develop some properties of the $a_{x,y}(\cdot)$-Neumann derivative for the nonlocal $s(\cdot,\cdot)$-order operator in fractional Musielak-Sobolev spaces with variable $s(\cdot,\cdot)-$order. Therefore we prove the basic proprieties of the correspondent function spaces. In the second part of this paper, by means of Ekeland's variational pri

  86. G. Valle, M. Dell'Omodarme, P. G. Prada Moroni, S. Degl'Innocenti

    A recent investigation highlighted peculiar trends between the radii derived from surface brightness-colour relations (SBCRs) combined with Gaia DR3 parallaxes with respect to asteroseismic scaling relation radii from K2 data. [...] We investigated on the robustness of the results based on Kepler data. We cross-matched asteroseismic and astrometric data for

  87. Jiale Cheng, Xiao Liu, Cunxiang Wang, Xiaotao Gu

    Instruction-following is a fundamental capability of language models, requiring the model to recognize even the most subtle requirements in the instructions and accurately reflect them in its output. Such an ability is well-suited for and often optimized by preference learning. However, existing methods often directly sample multiple independent responses fr

  88. A. A. Gerasimov, D. R. Lebedev, S. V. Oblezin

    The $GL_{\ell+1}(\mathbb{R})$ Hecke-Baxter operator was introduced as an element of the $O_{\ell+1}$-spherical Hecke algebra associated with the Gelfand pair $O_{\ell+1}\subset GL_{\ell+1}(\mathbb{R})$. It was specified by the property to act on an $O_{\ell+1}$-fixed vector in a $GL_{\ell+1}(\mathbb{R})$-principal series representation via multiplication by

  89. Wen-Chun Chen, Santiago Berrezueta-Guzman, Stefan Wagner

    Intellectual Disability (ID) is characterized by deficits in intellectual functioning and adaptive behavior, necessitating customized therapeutic interventions to improve daily life skills. This paper presents the development and evaluation of Space Exodus, a task-based role-playing Virtual Reality (VR) game designed to support therapy for children with ID.

  90. Anton J. Heckens, Efstratios Manolakis, Cedric Schuhmann, Thomas Guhr

    Multivariate Distributions are needed to capture the correlation structure of complex systems. In previous works, we developed a Random Matrix Model for such correlated multivariate joint probability density functions that accounts for the non-stationarity typically found in complex systems. Here, we apply these results to the returns measured in correlated

  91. Efstratios Manolakis, Anton J. Heckens, Benjamin Köhler, Thomas Guhr

    Risk assessment for rare events is essential for understanding systemic stability in complex systems. As rare events are typically highly correlated, it is important to study heavy-tailed multivariate distributions of the relevant variables, especially in the presence of non-stationarity. We use a generalized scalar product between correlation matrices to cl

  92. Huhu Zhang, Xing Gao

    Rota-Baxter operators on groups were studied quite recently. Motivated mainly by the fact that weight zero Rota-Baxter operators and averaging operators are Koszul dual to each other, we propose the concepts of averaging group and averaging Hopf algebra, and study relationships among them and the existing averaging Lie algebras. We also show that an averagin

  93. Zichen Tang, Hongyu Yang, Hanchen Zhang, Jiaxin Chen

    Advancements in neural implicit representations and differentiable rendering have markedly improved the ability to learn animatable 3D avatars from sparse multi-view RGB videos. However, current methods that map observation space to canonical space often face challenges in capturing pose-dependent details and generalizing to novel poses. While diffusion mode

  94. Lorenzo Carlucci, Oriola Gjetaj, Quentin Le Houérou, Ludovic Levy Patey

    The family of finite subsets $s$ of the natural numbers such that $|s|=1+\min s$ is known as the Schreier barrier in combinatorics and Banach Space theory, and as the family of exactly $\omega$-large sets in Logic. We formulate and prove the generalizations of Friedman's Free Set and Thin Set theorems and of Rainbow Ramsey's theorem to colorings of the Schre

  95. Anton Pichler

    Despite dramatic growth and cost improvements in renewables, existing energy companies exhibit significant inertia in adapting to the evolving technological landscape. This study examines technology transition patterns by analyzing over 140,000 investments in power assets over more than two decades, focusing on how firms expand existing technology holdings a

  96. Daoyi Gao, Yawar Siddiqui, Lei Li, Angela Dai

    Articulated 3D object generation is fundamental for creating realistic, functional, and interactable virtual assets which are not simply static. We introduce MeshArt, a hierarchical transformer-based approach to generate articulated 3D meshes with clean, compact geometry, reminiscent of human-crafted 3D models. We approach articulated mesh generation in a pa

  97. Graeme D. Berk, Simon Milz, Kavan Modi

    In arXiv:2110.02613, we presented a generalised dynamical resource theory framework that enabled noise reduction techniques including dynamical decoupling (DD) to be studied. While this fundamental contribution remains correct, it has been found that the main resource quantifiers we employed to study these resource theories -- based on the relative entropies

  98. Marjolaine Ray, Qi Wang, Frédérique Mélanie-Becquet, Thierry Poibeau

    Event detection in text streams is a crucial task for the analysis of online media and social networks. One of the current challenges in this field is establishing a performance standard while maintaining an acceptable level of computational complexity. In our study, we use an incremental clustering algorithm combined with recent advancements in sentence emb

  99. Zhipeng Chen, Lan Yang, Yonggang Qi, Honggang Zhang

    Despite the rapid advancements in text-to-image (T2I) synthesis, enabling precise visual control remains a significant challenge. Existing works attempted to incorporate multi-facet controls (text and sketch), aiming to enhance the creative control over generated images. However, our pilot study reveals that the expressive power of humans far surpasses the c

  100. Matthias Knopf, Sandra Barna, Daniel Radmanovac, Thomas Bergauer

    Microdosimetry investigates the energy deposition of ionizing radiation at microscopic scales, beyond the assessment capabilities of macroscopic dosimetry. This contributes to an understanding of the biological response in radiobiology, radiation protection and radiotherapy. Microdosimetric pulse height spectra are usually measured using an ionization detect