Skip to content

October 2024 arXiv papers — page 93

Showing 9,2019,300 of 23,665 papers

  1. Ali Eghbali

    We proceed to construct a dual pair for the $AdS_3 \times S^3$ background by applying non-Abelian T-duality (here as Poisson-Lie (PL) T-duality on a semi-Abelian double). By using a certain parametrization of the $4$-dimensional Lie group ${A}_2 \otimes 2{A}_1$ and by a suitable choice of spectator-dependent matrices, the original $\sigma$-model including th

  2. Zihui Wu, Haichang Gao, Ping Wang, Shudong Zhang

    Glitch tokens, inputs that trigger unpredictable or anomalous behavior in Large Language Models (LLMs), pose significant challenges to model reliability and safety. Existing detection methods primarily rely on heuristic embedding patterns or statistical anomalies within internal representations, limiting their generalizability across different model architec

  3. Fengjun Pan, Xiaobao Wu, Zongrui Li, Anh Tuan Luu

    Fallacies are defective arguments with faulty reasoning. Detecting and classifying them is a crucial NLP task to prevent misinformation, manipulative claims, and biased decisions. However, existing fallacy classifiers are limited by the requirement for sufficient labeled data for training, which hinders their out-of-distribution (OOD) generalization abilitie

  4. Jiaxin Yu, Tianchen Qian

    To optimize mobile health interventions and advance domain knowledge on intervention design, it is critical to understand how the intervention effect varies over time and with contextual information. This study aims to assess how a push notification suggesting physical activity influences individuals' step counts using data from the HeartSteps micro-randomiz

  5. Siyuan Lu, Jiaqi Shao, Bing Luo, Tao Lin

    Large Language Model (LLM) based multi-agent systems (MAS) have shown promise in tackling complex tasks, but often rely on predefined roles and centralized coordination, limiting their adaptability to evolving challenges. This paper introduces MorphAgent, a novel Autonomous, Self-Organizing, and Self-Adaptive Multi-Agent System for decentralized agent collab

  6. Tugrul Cabir Hakyemez, Omer Adar

    Accurate forecasting of electrical demand is essential for maintaining a stable and reliable power grid, optimizing the allocation of energy resources, and promoting efficient energy consumption practices. This study investigates the effectiveness of five hyperparameter optimization (HPO) algorithms -- Random Search, Covariance Matrix Adaptation Evolution St

  7. Huihui Yang, Chunxue Zhu, Longlong Lin, Pingpeng Yuan

    Identifying communities from temporal networks facilitates the understanding of potential dynamic relationships among entities, which has already received extensive applications. However, existing methods primarily rely on lower-order connectivity (e.g., temporal edges) to capture the structural and temporal cohesiveness of the community, often neglecting hi

  8. Jiaqi Shao, Tao Lin, Xiaojin Zhang, Qiang Yang

    Federated Unlearning (FU) enables the removal of specific clients' data influence from trained models. However, in non-IID settings, removing clients creates critical side effects: remaining clients with similar data distributions suffer disproportionate performance degradation, while the global model's stability deteriorates. These vulnerable clients then h

  9. Shuning Zhang, Xin Yi, Haobin Xing, Lyumanshan Ye

    Current Large Language Models (LLMs) cannot support users to precisely balance privacy protection and output performance during individual consultations. We introduce Adanonymizer, an anonymization plug-in that allows users to control this balance by navigating a trade-off curve. A survey (N=221) revealed a privacy paradox, where users frequently disclosed s

  10. Effie Papageorgiou

    Let $\mu$ be a radial compactly supported distribution on a harmonic $NA$ group. We prove that the right convolution operator $c_{\mu}:f \mapsto f* \mu$ maps the space of smooth $\mathfrak{v}$-radial functions onto itself if and only if the spherical Fourier transform $\widetilde{\mu}(\lambda)$, $\lambda \in \mathbb{C}$, is slowly decreasing. As an applicati

  11. Mengnan Zhao, Lihe Zhang, Jingwen Ye, Huchuan Lu

    Adversarial training (AT) refers to integrating adversarial examples -- inputs altered with imperceptible perturbations that can significantly impact model predictions -- into the training process. Recent studies have demonstrated the effectiveness of AT in improving the robustness of deep neural networks against diverse adversarial attacks. However, a compr

  12. Tian-Ming Li, Jia-Chi Zhang, Bing-Jie Chen, Kaixuan Huang

    For superconducting quantum processors, stable high-fidelity two-qubit operations depend on precise flux control of the tunable coupler. However, the pulse distortion poses a significant challenge to the control precision. Current calibration methods, which often rely on microwave crosstalk or additional readout resonators for coupler excitation and readout,

  13. Zichen Wang, Yaokun Ji, Jianing Tian, Shuangjia Zheng

    Antibodies are essential proteins responsible for immune responses in organisms, capable of specifically recognizing antigen molecules of pathogens. Recent advances in generative models have significantly enhanced rational antibody design. However, existing methods mainly create antibodies from scratch without template constraints, leading to model optimizat

  14. Yuedan Ding, Shidi Zhang, Henggeng Han, Wenyuan Cui

    Using the LAMOST DR7 low-resolution spectra, we carried out a systematic study of stellar chromospheric activity in both single and binary stars. We constructed a binary sample and a single-star sample, mainly using the binary belt and the main sequence in the Hertzsprung-Russell diagram, respectively. By comparing the $S$ indices between single and binary s

  15. Siyuan Yan, Zhen Yu, Clare Primiero, Cristina Vico-Alonso

    Diagnosing and treating skin diseases require advanced visual skills across domains and the ability to synthesize information from multiple imaging modalities. While current deep learning models excel at specific tasks like skin cancer diagnosis from dermoscopic images, they struggle to meet the complex, multimodal requirements of clinical practice. Here, we

  16. Thang Pang Ern

    We wish to discuss positive integer solutions to the Diophantine equation $$\prod_{k=1}^n(k^2+1)=b^2.$$ Some methods in analytic number theory will be used to tackle this problem.

  17. Nishat Raihan, Antonios Anastasopoulos, Marcos Zampieri

    Recent advancements in large language models (LLMs) have significantly enhanced code generation from natural language prompts. The HumanEval Benchmark, developed by OpenAI, remains the most widely used code generation benchmark. However, this and other Code LLM benchmarks face critical limitations, particularly in task diversity, test coverage, and linguisti

  18. Chenguang Lu

    Recent advances in deep learning suggest that we need to maximize and minimize two different kinds of information simultaneously. The Information Max-Min (IMM) method has been used in deep learning, reinforcement learning, and maximum entropy control. Shannon's information rate-distortion function is the theoretical basis of Minimizing Mutual Information (MM

  19. Xin Li, Wenhui Zhu, Xuanzhao Dong, Oana M. Dumitrascu

    With the rapid development of deep learning, CNN-based U-shaped networks have succeeded in medical image segmentation and are widely applied for various tasks. However, their limitations in capturing global features hinder their performance in complex segmentation tasks. The rise of Vision Transformer (ViT) has effectively compensated for this deficiency of

  20. Mingxin Li, Zhijie Nie, Yanzhao Zhang, Dingkun Long

    Text embeddings are vital for tasks such as text retrieval and semantic textual similarity (STS). Recently, the advent of pretrained language models, along with unified benchmarks like the Massive Text Embedding Benchmark (MTEB), has facilitated the development of versatile general-purpose text embedding models. Advanced embedding models are typically develo

  21. Haoran Feng, Zhiwei Chen, Zhibo Jiang, Yuehui Ma

    Structures in molecular ISM are observed to follow a power-law relation between the velocity dispersion and spatial size, known as Larson's first relation, which is often attributed to the turbulent nature of molecular ISM and imprints the dynamics of molecular cloud structures. Using the ${}^{13}\mathrm{CO}~(J=1-0)$ data from the Milky Way Imaging Scroll Pa

  22. Mingyi Zhou, Xiang Gao, Xiao Chen, Chunyang Chen

    Deploying DL models on mobile Apps has become ever-more popular. However, existing studies show attackers can easily reverse-engineer mobile DL models in Apps to steal intellectual property or generate effective attacks. A recent approach, Model Obfuscation, has been proposed to defend against such reverse engineering by obfuscating DL model representations,

  23. Sudipta Das, Ayan Patra, Rivu Gupta, Aditi Sen De

    In order to enable the sequential implementation of quantum information theoretic protocols in the continuous variable framework, we propose two schemes for resource reusability, resource-splitting protocol and unsharp homodyne measurements. We demonstrate the advantage offered by the first scheme in implementing sequential attempts at continuous variable te

  24. Zhiguang Zhang, Yuxiang Li

    In this work, we study the doubly degenerate nutrient taxis system with logistic source \begin{align} \begin{cases}\tag{$\star$}\label{eq 0.1} u_t=\nabla \cdot(u^{l-1} v \nabla u)- \nabla \cdot\left(u^{l} v \nabla v\right)+ u - u^2, \\ v_t=\Delta v-u v \end{cases} \end{align} in a smooth bounded domain $\Omega \subset \mathbb{R}^2$, where $l \geqslant 1$. It

  25. Thomas Bläsius, Max Göttlicher, Sascha Gritzbach, Wendy Yi

    Motivated by the cabling of solar farms, we study the problem Constrained Layer Tree. At its core, it asks whether there exists a tree that connects a set of sources (the leaves) to one sink (the root) such that certain capacity constraints at the inner nodes are satisfied. Our main algorithmic contribution is a dynamic program with various optimizations for

  26. Amelia Jones

    This research delves into the development of a fatigue detection system based on modern object detection algorithms, particularly YOLO (You Only Look Once) models, including YOLOv5, YOLOv6, YOLOv7, and YOLOv8. By comparing the performance of these models, we evaluate their effectiveness in real-time detection of fatigue-related behavior in drivers. The study

  27. Yuzhe Weng, Haotian Wang, Tian Gao, Kewei Li

    In multimodal sentiment analysis, collecting text data is often more challenging than video or audio due to higher annotation costs and inconsistent automatic speech recognition (ASR) quality. To address this challenge, our study has developed a robust model that effectively integrates multimodal sentiment information, even in the absence of text modality. S

  28. Dipo Dunsin, Mohamed Chahine Ghanem, Karim Ouazzane, Vassil Vassilev

    This Research proposes a Novel Reinforcement Learning (RL) model to optimise malware forensics investigation during cyber incident response. It aims to improve forensic investigation efficiency by reducing false negatives and adapting current practices to evolving malware signatures. The proposed RL framework leverages techniques such as Q-learning and the M

  29. Lianghua Huang, Wei Wang, Zhi-Fan Wu, Huanzhang Dou

    While large language models (LLMs) have revolutionized natural language processing with their task-agnostic capabilities, visual generation tasks such as image translation, style transfer, and character customization still rely heavily on supervised, task-specific datasets. In this work, we introduce Group Diffusion Transformers (GDTs), a novel framework tha

  30. Wenyi Liu, Rui Wang, Yuanshuai Luo, Jianjun Wei

    With the explosive growth of Internet data, users are facing the problem of information overload, which makes it a challenge to efficiently obtain the required resources. Recommendation systems have emerged in this context. By filtering massive amounts of information, they provide users with content that meets their needs, playing a key role in scenarios suc

  31. Minsun Kim, SeonGyeom Kim, Suyoun Lee, Yoosang Yoon

    This paper presents the development of a dashboard designed specifically for teachers in English as a Foreign Language (EFL) writing education. Leveraging LLMs, the dashboard facilitates the analysis of student interactions with an essay writing system, which integrates ChatGPT for real-time feedback. The dashboard aids teachers in monitoring student behavio

  32. Behnaz Omoomi Marzieh Vahid Dastjerdi

    The star chromatic index of a graph $G$, denoted by $\chi^\prime_s(G)$, is the smallest integer $k$ for which $G$ admits a proper edge coloring with $k$ colors such that every path and cycle of length four is not bicolored. Let $d$ be the greatest common divisor of $n$ and $k$. Zhu~et~al. (\footnotesize{Discussiones Mathematicae: Graph Theory, 41(2): 1265, 2

  33. Yuchi Yahagi, Rintaro Chujo, Yuga Harada, Changyo Han

    Listening to audio content, such as podcasts and audiobooks, is one way for people to engage with knowledge. Listening affords people more mobility than reading by seeing, thereby broadening their learning opportunities. This study explores the potential applications of large language models (LLMs) to adapt text documents to audio content and addresses the l

  34. Nguyen Thang Loi, Duong Tan Loc, Vo Nguyen Le Duy

    Feature Selection (FS) under domain adaptation (DA) is a critical task in machine learning, especially when dealing with limited target data. However, existing methods lack the capability to guarantee the reliability of FS under DA. In this paper, we introduce a novel statistical method to statistically test FS reliability under DA, named SFS-DA (statistical

  35. Hidetaka Kamigaito, Hiroyuki Deguchi, Yusuke Sakai, Katsuhiko Hayashi

    Inference methods play an important role in eliciting the performance of large language models (LLMs). Currently, LLMs use inference methods utilizing generated multiple samples, which can be derived from Minimum Bayes Risk (MBR) Decoding. Previous studies have conducted empirical analyses to clarify the improvements in generation performance achieved by MBR

  36. Baojian Zhou, Yifan Sun, Reza Babanezhad Harikandeh, Xingzhi Guo

    Given the damping factor $\alpha$ and precision tolerance $\epsilon$, \citet{andersen2006local} introduced Approximate Personalized PageRank (APPR), the \textit{de facto local method} for approximating the PPR vector, with runtime bounded by $\Theta(1/(\alpha\epsilon))$ independent of the graph size. Recently, \citet{fountoulakis2022open} asked whether faste

  37. Jinggui Liang, Yuxia Wu, Yuan Fang, Hao Fei

    In the rapidly evolving field of conversational AI, Ontology Expansion (OnExp) is crucial for enhancing the adaptability and robustness of conversational agents. Traditional models rely on static, predefined ontologies, limiting their ability to handle new and unforeseen user needs. This survey paper provides a comprehensive review of the state-of-the-art te

  38. Arturo Mariani, Federico Senocrate, Jason Mikiel-Hunter, David McAlpine

    Background: In Kreuz et al., J Neurosci Methods 381, 109703 (2022) two methods were proposed that perform latency correction, i.e., optimize the spike time alignment of sparse neuronal spike trains with well defined global spiking events. The first one based on direct shifts is fast but uses only partial latency information, while the other one makes use of

  39. Md Mubtasim Ahasan, Md Fahim, Tasnim Mohiuddin, A K M Mahbubur Rahman

    Recent advancements in speech-language models have yielded significant improvements in speech tokenization and synthesis. However, effectively mapping the complex, multidimensional attributes of speech into discrete tokens remains challenging. This process demands acoustic, semantic, and contextual information for precise speech representations. Existing spe

  40. Jiahao Wang, Amer Shalaby

    Users of the transit system flood social networks daily with messages that contain valuable insights crucial for improving service quality. These posts help transit agencies quickly identify emerging issues. Parsing topics and sentiments is key to gaining comprehensive insights to foster service excellence. However, the volume of messages makes manual analys

  41. Yue Zhan, Zhihong Zeng, Haijun Liu, Xiaoheng Tan

    The purpose of RGB-D Salient Object Detection (SOD) is to pinpoint the most visually conspicuous areas within images accurately. While conventional deep models heavily rely on CNN extractors and overlook the long-range contextual dependencies, subsequent transformer-based models have addressed the issue to some extent but introduce high computational complex

  42. Oleg Zubelevich

    In this short note we present a far generalization of the following very well-known assertion: assume that we have two orthonormal sequences in a Hilbert space and these sequences are quadratically close to each other. Then if one of these sequences is a basis in the Hilbert space then so is the other one.

  43. Weiyong He, Long Li, Xiaowei Xu

    The purpose of this article is to study the (residual) Monge-Amp\`{e}re mass of a plurisubharmonic function with an isolated unbounded locus. A general decomposition formula is obtained under the Sasakian structure of the unit sphere. In complex dimension two, we obtain an $L^{1}$-apriori estimate on the complex Monge-Amp\`{e}re operator. This induces an upp

  44. Jiahao Wang, Amer Shalaby

    Accurate prediction of public transit ridership is vital for efficient planning and management of transit in rapidly growing urban areas in Canada. Unexpected increases in passengers can cause overcrowded vehicles, longer boarding times, and service disruptions. Traditional time series models like ARIMA and SARIMA face limitations, particularly in short-term

  45. Gesa Mittmann, Sara Laiouar-Pedari, Hendrik A. Mehrtens, Sarah Haggenmüller

    The aggressiveness of prostate cancer, the most common cancer in men worldwide, is primarily assessed based on histopathological data using the Gleason scoring system. While artificial intelligence (AI) has shown promise in accurately predicting Gleason scores, these predictions often lack inherent explainability, potentially leading to distrust in human-mac

  46. Arnab Bhattacharya, Afsar Ahmed, Sreeparvathy PC, Daichi Kurebayashi

    The synergy between real and reciprocal space topology is anticipated to yield a diverse array of topological properties in quantum materials. We address this pursuit by achieving topologically safeguarded magnetic order in novel Weyl metallic Heusler alloy, Mn$_{2}$Pd$_{0.5}$Ir$_{0.5}$Sn. The system possesses non-centrosymmetric D$_{2d}$ crystal symmetry wi

  47. Sizhe Liu, Jun Xia, Lecheng Zhang, Yuchen Liu

    Molecular relational learning (MRL) is crucial for understanding the interaction behaviors between molecular pairs, a critical aspect of drug discovery and development. However, the large feasible model space of MRL poses significant challenges to benchmarking, and existing MRL frameworks face limitations in flexibility and scope. To address these challenges

  48. M. Rostami, S. S. Kia

    In this article, we consider the problem of unconstrained time-varying convex optimization, where the cost function changes with time. We provide an in-depth technical analysis of the problem and argue why freezing the cost at each time step and taking finite steps toward the minimizer is not the best tracking solution for this problem. We propose a set of a

  49. Minseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon

    Accelerating end-to-end inference of transformer-based large language models (LLMs) is a critical component of AI services in datacenters. However, diverse compute characteristics of end-to-end LLM inference present challenges as previously proposed accelerators only address certain operations or stages (e.g., self-attention, generation stage, etc.). To addr

  50. Ying Hu, Chenyi Zhuang, Pan Gao

    Style transfer aims to fuse the artistic representation of a style image with the structural information of a content image. Existing methods train specific networks or utilize pre-trained models to learn content and style features. However, they rely solely on textual or spatial representations that are inadequate to achieve the balance between content and

  51. Baohua Huang, Jiakai Chen, Wen Li

    One of the tasks in color image processing and computer vision is to recover clean data from partial observations corrupted by noise. To this end, robust quaternion matrix completion (QMC) has recently attracted more attention and shown its effectiveness, whose convex relaxation is to minimize the quaternion nuclear norm plus the quaternion $L_1$-norm. Howev

  52. Zhiqiang Zhong, Yang Yang, Fengqiang Wan, Henglu Wei

    This report presents our method for Single Object Tracking (SOT), which aims to track a specified object throughout a video sequence. We employ the LoRAT method. The essence of the work lies in adapting LoRA, a technique that fine-tunes a small subset of model parameters without adding inference latency, to the domain of visual tracking. We train our model u

  53. Yi Zhao, Jing Li, Linyi Yang

    Large language models (LLMs) are widely used, but concerns about data contamination challenge the reliability of LLM evaluations. Existing contamination detection methods are often task-specific or require extra prerequisites, limiting practicality. We propose a novel framework, Consistency Amplification-based Data Contamination Detection (CAP), which introd

  54. Robert Mercaş, Wen Chean Teh

    The focus of this work is the study of Parikh matrices with emphasis on two concrete problems. In the first part of our presentation we show that a conjecture by Dick at al. in 2021 only stands in the case of ternary alphabets, while providing counterexamples for larger alphabets. In particular, we show that the only type of distinguishability in the case of

  55. Chen Yan, Weina Wang, Lei Ying

    We study the finite-horizon Restless Multi-Armed Bandit (RMAB) problem with $N$ homogeneous arms. Prior work has shown that when an RMAB satisfies a non-degeneracy condition, Linear-Programming-based (LP-based) policies derived from the fluid approximation, which captures the mean dynamics of the system, achieve an exponentially small optimality gap. However

  56. Sahil Verma, Royi Rassin, Arnav Das, Gantavya Bhatt

    Text-to-image models are trained using large datasets of image-text pairs collected from the internet. These datasets often include copyrighted and private images. Training models on such datasets enables them to generate images that might violate copyright laws and individual privacy. This phenomenon is termed imitation -- generation of images with content

  57. Shubhajit Roy, Hrriday Ruparel, Kishan Ved, Anirban Dasgupta

    Scalability of Graph Neural Networks (GNNs) remains a significant challenge. To tackle this, methods like coarsening, condensation, and computation trees are used to train on a smaller graph, resulting in faster computation. Nonetheless, prior research has not adequately addressed the computational costs during the inference phase. This paper presents a nove

  58. Songnian Xu, Wenhao Zhen, Dein Wong

    Let $G$ be a simple connected graph of order $n$ with diameter $d$. Let $m_G(-1)$ denote the multiplicity of the eigenvalue $-1$ of the adjacency matrix of $G$, and let $P = P_{d+1}$ be the diameter path of $G$. If $-1$ is not an eigenvalue of $P$, then by the interlacing theorem, we have $m_G(-1)\leq n - d - 1$. In this article, we characterize the extremal

  59. Peter Kuchment, Leonid Kunyansky

    The forward problem arising in several hybrid imaging modalities can be modeled by the Cauchy problem for the free space wave equation. Solution to this problems describes propagation of a pressure wave, generated by a source supported inside unit sphere $S$. The data $g$ represent the time-dependent values of the pressure on the observation surface $S$. Fin

  60. Raymundo Vazquez Martinez, Raj Abhijit Dandekar, Rajat Dandekar, Sreedath Panat

    In this study, we apply two pillars of Scientific Machine Learning: Neural Ordinary Differential Equations (Neural ODEs) and Universal Differential Equations (UDEs) to the Chandrasekhar White Dwarf Equation (CWDE). The CWDE is fundamental for understanding the life cycle of a star, and describes the relationship between the density of the white dwarf and its

  61. Tuan Nam Nguyen, Seymanur Akti, Ngoc Quan Pham, Alexander Waibel

    Previous approaches on accent conversion (AC) mainly aimed at making non-native speech sound more native while maintaining the original content and speaker identity. However, non-native speakers sometimes have pronunciation issues, which can make it difficult for listeners to understand them. Hence, we developed a new AC approach that not only focuses on acc

  62. Ankur Mudgal, Abhishek Verma, Munesh Singh, Kshira Sagar Sahoo

    Software Defined Networking (SDN) has evolved to revolutionize next-generation networks, offering programmability for on-the-fly service provisioning, primarily supported by the OpenFlow (OF) protocol. The limited storage capacity of Ternary Content Addressable Memory (TCAM) for storing flow tables in OF switches introduces vulnerabilities, notably the Low-R

  63. Junkai Jiang, Zeyu Han, Yuning Wang, Mengchi Cai

    Driving risk assessment is crucial for both autonomous vehicles and human-driven vehicles. The driving risk can be quantified as the product of the probability that an event (such as collision) will occur and the consequence of that event. However, the probability of events occurring is often difficult to predict due to the uncertainty of drivers' or vehicle

  64. Michał Borowski, Pierre Bousquet, Iwona Chlebicka, Benjamin Lledos

    We establish that the Lavrentiev gap between Sobolev and Lipschitz maps does not occur for a scalar variational problem of the form: \[ \textrm{to minimize} \qquad u \mapsto \int_\Omega f(x,u,\nabla u)\,dx \,, \] under a Dirichlet boundary condition. Here, \(\Omega\) is a bounded Lipschitz open set in \(\rn\), \(N\geq 1\) and the function $f$ is required to

  65. Prateek Chennuri, Yiheng Chi, Enze Jiang, G. M. Dilshan Godaliyadda

    The proliferation of single-photon image sensors has opened the door to a plethora of high-speed and low-light imaging applications. However, data collected by these sensors are often 1-bit or few-bit, and corrupted by noise and strong motion. Conventional video restoration methods are not designed to handle this situation, while specialized quanta burst alg

  66. Hao Wu, Donglin Bai, Shiqi Jiang, Qianxi Zhang

    Video activity recognition has become increasingly important in robots and embodied AI. Recognizing continuous video activities poses considerable challenges due to the fast expansion of streaming video, which contains multi-scale and untrimmed activities. We introduce a novel system, CARS, to overcome these issues through adaptive video context modeling. Ad

  67. Woojin Chae, Kihyuk Hong, Yufan Zhang, Ambuj Tewari

    This paper proposes a computationally tractable algorithm for learning infinite-horizon average-reward linear mixture Markov decision processes (MDPs) under the Bellman optimality condition. Our algorithm for linear mixture MDPs achieves a nearly minimax optimal regret upper bound of $\widetilde{\mathcal{O}}(d\sqrt{\mathrm{sp}(v^*)T})$ over $T$ time steps wh

  68. Deeparghya Dutta Barua, Md Sakib Ul Rahman Sourove, Md Fahim, Fabiha Haider

    Visual Question Answer (VQA) poses the problem of answering a natural language question about a visual context. Bangla, despite being a widely spoken language, is considered low-resource in the realm of VQA due to the lack of proper benchmarks, challenging models known to be performant in other languages. Furthermore, existing Bangla VQA datasets offer littl

  69. Sivangi Chatterjee, Srishti Ganguly, Avik Bose, Hrithik Raj Prasad

    This project explores the application of machine learning techniques for music genre classification using the GTZAN dataset, which contains 100 audio files per genre. Motivated by the growing demand for personalized music recommendations, we focused on classifying five genres-Blues, Classical, Jazz, Hip Hop, and Country-using a variety of algorithms includin

  70. Longtao Zhu, Hongyu Yang, Ge Song, Xin Ma

    Current flight procedure design methods heavily rely on human-led design process, which is not only low auto-mation but also suffer from complex algorithm modelling and poor generalization. To address these challenges, this paper proposes an agent-driven flight procedure design method based on large language model, named Au-toFPDesigner, which utilizes multi

  71. Noam Ginio, Michael Lindenbaum, Barak Fishbain, Dan Liberzon

    Effective spatio-temporal measurements of water surface elevation (water waves) in laboratory experiments are crucial for scientific and engineering research. Existing techniques are often cumbersome, computationally heavy and generally suffer from limitations in wavenumber/frequency response. To address these challenges, we propose Wave (from) Polarized Lig

  72. Zhewei Dai, Shilei Zeng, Haotian Liu, Xurui Li

    We introduce SeaS, a unified industrial generative model for automatically creating diverse anomalies, authentic normal products, and precise anomaly masks. While extensive research exists, most efforts either focus on specific tasks, i.e., anomalies or normal products only, or require separate models for each anomaly type. Consequently, prior methods either

  73. Yunqi Cai, Jiangnan Li, Dong Wang

    Micromagnetics has made significant strides, particularly due to its wide-ranging applications in magnetic storage design. Numerical simulation is a cornerstone of micromagnetics research, relying on first-principle rules to compute the dynamic evolution of micromagnetic systems based on the renowned LLG equation, named after Landau, Lifshitz, and Gilbert. H

  74. Andrew Fleck, Edward Furman, Yang Shen

    Nowadays insurers have to account for potentially complex dependence between risks. In the field of loss reserving, there are many parametric and non-parametric models attempting to capture dependence between business lines. One common approach has been to use additive background risk models (ABRMs) which provide rich and interpretable dependence structures

  75. Andrew Fleck, Edward Furman, Yang Shen

    In order to properly manage risk, practitioners must understand the aggregate risks they are exposed to. Additionally, to properly price policies and calculate bonuses the relative riskiness of individual business units must be well understood. Certainly, Insurers and Financiers are interested in the properties of the sums of the risks they are exposed to an

  76. Huyen Le, Khiet Dang, Nhung Nguyen, Mai Tran

    Human-induced pluripotent stem cell-derived cardiomyocytes (hiPSC-CMs) are a powerful tool in advancing cardiovascular research and clinical applications. The maturation of sarcomere organization in hiPSC-CMs is crucial, as it supports the contractile function and structural integrity of these cells. Traditional methods for assessing this maturation like man

  77. Vindhyawasini Prasad

    Dark matter (DM) is a new type of invisible matter introduced to explain various features of recent astrophysical observations, including galaxy rotation curves and other fundamental characteristics of our universe. DM may couple to ordinary matter via portals, which open up possibilities for new particles, such as axion-like particle, light Higgs boson, dar

  78. William Zhao

    We introduce the stack-sorting map $\text{SC}_\sigma$ that sorts, in a right-greedy manner, an input permutation through a stack that avoids some vincular pattern $\sigma$. The stack-sorting maps of Cerbai et al. in which the stack avoids a pattern classically and Defant and Zheng in which the stack avoids a pattern consecutively follow as special cases. We

  79. Ken Sekimoto

    In the previous paper we have shown analytically that, if the drift function of the d-dimensional Langevin equation is the Langevin function with a properly chosen scale factor, then the evolution of the drift function is a martingale associated with the histories generated by the very Langevin equation. Moreover, we numerically demonstrated that those gener

  80. Kun Wang, Zhiqiang Yan, Junkai Fan, Wanlu Zhu

    In this paper, we introduce DCDepth, a novel framework for the long-standing monocular depth estimation task. Moving beyond conventional pixel-wise depth estimation in the spatial domain, our approach estimates the frequency coefficients of depth patches after transforming them into the discrete cosine domain. This unique formulation allows for the modeling

  81. Wei Xie, Shuoyoucheng Ma, Zhenhua Wang, Enze Wang

    The cognitive mechanism by which Large Language Models (LLMs) solve mathematical problems remains a widely debated and unresolved issue. Currently, there is little interpretable experimental evidence that connects LLMs' problem-solving with human cognitive psychology.To determine if LLMs possess human-like mathematical reasoning, we modified the problems use

  82. Kent K. Chang, Anna Ho, David Bamman

    Television is often seen as a site for subcultural identification and subversive fantasy, including in queer cultures. How might we measure subversion, or the degree to which the depiction of social relationship between a dyad (e.g. two characters who are colleagues) deviates from its typical representation on TV? To explore this question, we introduce the t

  83. Linh Van Ma, Muhammad Ishfaq Hussain, Kin-Choong Yow, Moongu Jeon

    The MS-GLMB filter offers a robust framework for tracking multiple objects through the use of multi-sensor data. Building on this, the MV-GLMB and MV-GLMB-AB filters enhance the MS-GLMB capabilities by employing cameras for 3D multi-sensor multi-object tracking, effectively addressing occlusions. However, both filters depend on overlapping fields of view fro

  84. Woong Bae Jeon, Dong Hyun Park, Jong Sung Moon, Kyu-Young Kim

    Scalable, reliable quantum light sources are essential for increasing quantum channel capacity and advancing quantum protocols based on photonic qubits. Although recent developments in solid-state quantum emitters have enabled the generation of single photons with high performance, the scalable integration of multiple quantum light sources onto practical opt

  85. Silong Yong, Yaqi Xie, Simon Stepputtis, Katia Sycara

    Volume rendering in neural radiance fields is inherently time-consuming due to the large number of MLP calls on the points sampled per ray. Previous works would address this issue by introducing new neural networks or data structures. In this work, We propose GL-NeRF, a new perspective of computing volume rendering with the Gauss-Laguerre quadrature. GL-NeRF

  86. Jihyo Kim, Seulbi Lee, Sangheum Hwang

    With the recent emergence of foundation models trained on internet-scale data and demonstrating remarkable generalization capabilities, such foundation models have become more widely adopted, leading to an expanding range of application domains. Despite this rapid proliferation, the trustworthiness of foundation models remains underexplored. Specifically, th

  87. Shangning Xia, Hongjie Fang, Cewu Lu, Hao-Shu Fang

    Generalization in robotic manipulation remains a critical challenge, particularly when scaling to new environments with limited demonstrations. This paper introduces CAGE, a novel robotic manipulation policy designed to overcome these generalization barriers by integrating a causal attention mechanism. CAGE utilizes the powerful feature extraction capabiliti

  88. Takeshi Gotoda

    We investigate enstrophy variations by collapse of point vortices in an inviscid flow and, in particular, focus on the enstrophy dissipation that is a significant property characterizing 2D turbulent flows. Point vortex is an ideal vortex whose vorticity is concentrated on a point and the dynamics of point vortices on an inviscid flow is described by the poi

  89. Suning Huang, Zheyu Zhang, Tianhai Liang, Yihan Xu

    Visual deep reinforcement learning (RL) enables robots to acquire skills from visual input for unstructured tasks. However, current algorithms suffer from low sample efficiency, limiting their practical applicability. In this work, we present MENTOR, a method that improves both the architecture and optimization of RL agents. Specifically, MENTOR replaces the

  90. Jilong Li, Zhenxi Song, Jiaqi Wang, Meishan Zhang

    Current EEG/MEG-to-text decoding systems suffer from three key limitations: (1) reliance on teacher-forcing methods, which compromises robustness during inference, (2) sensitivity to session-specific noise, hindering generalization across subjects, and (3) misalignment between brain signals and linguistic representations due to pre-trained language model ove

  91. Xiaohang Xu, Renhe Jiang, Chuang Yang, Zipei Fan

    With the popularity of location-based services, human mobility prediction plays a key role in enhancing personalized navigation, optimizing recommendation systems, and facilitating urban mobility and planning. This involves predicting a user's next POI (point-of-interest) visit using their past visit history. However, the uneven distribution of visitations o

  92. Marie Roald, Magnus Breder Birkenes, Lars Gunnarsønn Bagøien Johnsen

    Digital tools for text analysis have long been essential for the searchability and accessibility of digitised library collections. Recent computer vision advances have introduced similar capabilities for visual materials, with deep learning-based embeddings showing promise for analysing visual heritage. Given that many books feature visuals in addition to te

  93. Ryan Diaz, Adam Imdieke, Vivek Veeriah, Karthik Desingh

    Operating in unstructured environments like households requires robotic policies that are robust to out-of-distribution conditions. Although much work has been done in evaluating robustness for visuomotor policies, the robustness evaluation of a multisensory approach that includes force-torque sensing remains largely unexplored. This work introduces a novel,

  94. Tarkes Dora Pallicity, Maximillian Krause, Thomas Böhlke

    Statistical fluctuations of local tensorial fields beyond the mean are relevant to predict localized failure or overall behavior of the inelastic composites. The expression for second moments of the local fields can be established using the Hill-Mandel condition. Complete estimation of statistical fluctuations via second moments is usually ignored despite it

  95. Naz Yaldiz, Amaresh Chakrabarti

    The design of an inclusive product lifecycle is important for empowering stakeholders through their meaningful inclusion in lifecycle processes. To achieve this, the inclusion of stakeholders must be structured in a way that supports their empowerment. Inclusivity addresses the lifecycle context to improve how diverse stakeholders are included across phases,

  96. Zilong Li

    This papers presents the submission of team Ryu to the canceled SIGMORPHON 2024 shared task on subword tokenization. My submission explores whether morphological segmentation methods can be used as a part of subword tokenizers. I adopt two approaches: the statistical segmentation method Morfessor and a transformer based sequence-to-sequence (seq2seq) segment

  97. Michael Huylo, Sina Taheri, Atila Novoselac

    There is currently a large federal effort to decarbonize the country's electrical grid as part of the clean energy transition. The elimination of fossil fuel fired systems, and their replacement with intermittent renewable sources and other electric equipment will require better load management techniques to ensure a reliable grid. One strategy for maintaini

  98. Haichuan Zhang, Meiyu Lin, Zhaoyi Liu, Renyuan Li

    As generative models achieve great success, tampering and modifying the sensitive image contents (i.e., human faces, artist signatures, commercial logos, etc.) have induced a significant threat with social impact. The backdoor attack is a method that implants vulnerabilities in a target model, which can be activated through a trigger. In this work, we innova

  99. Hongqiu Wang, Zhaohu Xing, Weitong Wu, Yijun Yang

    Fundus imaging is a pivotal tool in ophthalmology, and different imaging modalities are characterized by their specific advantages. For example, Fundus Fluorescein Angiography (FFA) uniquely provides detailed insights into retinal vascular dynamics and pathology, surpassing Color Fundus Photographs (CFP) in detecting microvascular abnormalities and perfusion

  100. Yuhao Gong, Yuchen Zhang, Fei Wang, Chi-Han Lee

    As global climate change intensifies, accurate weather forecasting has become increasingly important, affecting agriculture, energy management, environmental protection, and daily life. This study introduces a hybrid model combining Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks to predict historical temperature data. CNNs ar