Skip to content

October 2024 arXiv papers — page 101

Showing 10,00110,100 of 23,665 papers

  1. Yuliang Gu, Yepeng Liu, Zhichao Sun, Jinchi Zhu

    Annotating 3D medical images demands expert knowledge and is time-consuming. As a result, semi-supervised learning (SSL) approaches have gained significant interest in 3D medical image segmentation. The significant size differences among various organs in the human body lead to imbalanced class distribution, which is a major challenge in the real-world appli

  2. Zihan Liu, Ruinan Zeng, Dongxia Wang, Gengyun Peng

    In industrial control systems, the generation and verification of Programmable Logic Controller (PLC) code are critical for ensuring operational efficiency and safety. While Large Language Models (LLMs) have made strides in automated code generation, they often fall short in providing correctness guarantees and specialized support for PLC programming. To add

  3. Xiaochuan Li, Zichun Yu, Chenyan Xiong

    Synthetic data has been widely used to train large language models, but their generative nature inevitably introduces noisy, non-informative, and misleading learning signals. In this paper, we propose Montessori-Instruct, a novel data synthesis framework that tailors the data synthesis ability of the teacher language model toward the student language model's

  4. Mathieu Seraphim, Alexis Lechervy, Florian Yger, Luc Brun

    Purpose: In sleep medicine, assessing the evolution of a subject's sleep often involves the costly manual scoring of electroencephalographic (EEG) signals. In recent years, a number of Deep Learning approaches have been proposed to automate this process, mainly by extracting features from said signals. However, despite some promising developments in related

  5. Mushir Akhtar, A. Quadir, M. Tanveer, Mohd. Arshad

    Alzheimer's disease (AD) is a leading neurodegenerative condition and the primary cause of dementia, characterized by progressive cognitive decline and memory loss. Its progression, marked by shrinkage in the cerebral cortex, is irreversible. Numerous machine learning algorithms have been proposed for the early diagnosis of AD. However, they often struggle w

  6. Henrik Seckler, Ralf Metzler

    When recording the movement of individual animals, cells or molecules one will often observe changes in their diffusive behaviour at certain points in time along their trajectory. In order to capture the different diffusive modes assembled in such heterogeneous trajectories it becomes necessary to segment them by determining these change-points. Such a chang

  7. Xiyuan Zhang, Diyan Teng, Ranak Roy Chowdhury, Shuheng Li

    Motion time series collected from mobile and wearable devices such as smartphones and smartwatches offer significant insights into human behavioral patterns, with wide applications in healthcare, automation, IoT, and AR/XR due to their low-power, always-on nature. However, given security and privacy concerns, building large-scale motion time series datasets

  8. Maciej Skorski

    Security of oscillatory true random number generators remains not fully understood due to insufficient understanding of complex $1/f^\alpha$ phase noise. To bridge this gap, we introduce fractional Brownian motion as a comprehensive theoretical framework, capturing power-law spectral densities from white to flicker frequency noise. Our key contributions prov

  9. Vishal Vivek Saley, Goonjan Saha, Rocktim Jyoti Das, Dinesh Raghu

    Medical task-oriented dialogue systems can assist doctors by collecting patient medical history, aiding in diagnosis, or guiding treatment selection, thereby reducing doctor burnout and expanding access to medical services. However, doctor-patient dialogue datasets are not readily available, primarily due to privacy regulations. Moreover, existing datasets l

  10. Shuang Geng, Zelin Ning, Fu Zhang, Boyu Zhou

    Autonomous exploration is a fundamental problem for various applications of unmanned aerial vehicles (UAVs). Recently, LiDAR-based exploration has gained significant attention due to its ability to generate high-precision point cloud maps of large-scale environments. While the point clouds are inherently informative for navigation, many existing exploration

  11. SeongYeub Chu, JongWoo Kim, Bryan Wong, MunYong Yi

    Existing automated essay scoring (AES) has solely relied on essay text without using explanatory rationales for the scores, thereby forgoing an opportunity to capture the specific aspects evaluated by rubric indicators in a fine-grained manner. This paper introduces Rationale-based Multiple Trait Scoring (RMTS), a novel approach for multi-trait essay scoring

  12. Asma Yamani, Malak Baslyman

    Text-to-Image generative systems are progressing rapidly to be a source of advertisement and media and could soon serve as image searches or artists. However, there is a significant concern about the representativity bias these models embody and how these biases can propagate in the social fabric after fine-tuning them. Therefore, continuously monitoring and

  13. Haoran Lai, Zihang Jiang, Qingsong Yao, Rongsheng Wang

    The development of 3D medical vision-language models holds significant potential for disease diagnosis and patient treatment. However, compared to 2D medical images, 3D medical images, such as CT scans, face challenges related to limited training data and high dimension, which severely restrict the progress of 3D medical vision-language models. To address th

  14. Basile Coron

    Motivated by the question of whether Chow polynomials of matroids have only real roots, this article revisits the known relationship between Eulerian polynomials and the Hilbert series of Chow rings of permutohedral varieties. This is done using a quadratic Gr\"obner basis associated to a new presentation of those rings, which is obtained by iterating the se

  15. Xiang Zhang, Dujian Ding

    Large Language Models (LLMs) have revolutionized natural language processing and hold immense potential for advancing Artificial Intelligence. However, the core architecture of most mainstream LLMs -- the Transformer -- has inherent limitations in computational depth, rendering them theoretically incapable of solving many reasoning tasks that demand increasi

  16. Sujitha Sathiyamoorthy, N Mohana, Anusha Prakash, Hema A Murthy

    The performance of a text-to-speech (TTS) synthesis model depends on various factors, of which the quality of the training data is of utmost importance. Millions of data are collected around the globe for various languages, but resources for Indian languages are few. Although there are many efforts involved in data collection, a common set of protocols for d

  17. Yidian Li, Xian Du, Junjie Wang, Runzhe Xu

    The surface of three-dimensional materials provides an ideal and versatile platform to explore quantum-confined physics. Here, we systematically investigate the electronic structure of Na-intercalated CrTe2, a van der Waals antiferromagnet, using angle-resolved photoemission spectroscopy and ab-initio calculations. The measured band structure deviates from t

  18. Honglin Li, Yunlong Zhang, Pingyi Chen, Zhongyi Shui

    Histopathology Whole Slide Image (WSI) analysis serves as the gold standard for clinical cancer diagnosis in the daily routines of doctors. To develop computer-aided diagnosis model for WSIs, previous methods typically employ Multi-Instance Learning to enable slide-level prediction given only slide-level labels. Among these models, vanilla attention mechanis

  19. Edward Finkelstein

    This paper presents a novel method for solving the 2D advection-diffusion equation using fixed-depth symbolic regression and symbolic differentiation without expression trees. The method is applied to two cases with distinct initial and boundary conditions, demonstrating its accuracy and ability to find approximate solutions efficiently. This framework offer

  20. Masashi Takeshita, Rafal Rzepka

    Natural Language Processing (NLP) research on AI Safety and social bias in AI has focused on safety for humans and social bias against human minorities. However, some AI ethicists have argued that the moral significance of nonhuman animals has been ignored in AI research. Therefore, the purpose of this study is to investigate whether there is speciesism, i.e

  21. Sehun Kim

    A persistence diagram provides a compact summary of persistent homology, which captures the topological features of a space at different scales. However, due to its nature as a set, incorporating it as a feature into a machine learning framework is challenging. Several methods have been proposed to use persistence diagrams as input for machine learning model

  22. Ken Deng, Zhongchi Zhang, Huaichuan Wang, Zihan Zhao

    We study the spatially incoherent light generated by a multimode fiber(MMF) in the application of image projection designed for the ultracold-atom experiments. Inspired by previous half-analytic methods concerning the incoherent light, here a full-numerical model is established to provide more quantitative descriptions, and part of results is compared with e

  23. Juyan Zhang, Dana Kulic, Michael Burke

    Manipulation tasks often consist of subtasks, each representing a distinct skill. Mastering these skills is essential for robots, as it enhances their autonomy, efficiency, adaptability, and ability to work in their environment. Learning from demonstrations allows robots to rapidly acquire new skills without starting from scratch, with demonstrations typical

  24. George E. Andrews, Mohamed El Bachraoui

    For a fixed positive integer $k$, let $C(k,n)$ denote the number of two-color partitions of $n$ with odd smallest part and restrictions on even parts, and let $C_k(q)$ be its generating function. We show that $C(1,n)\equiv d(2n-1)\pmod{4}$ and obtain congruences modulo $2$ and $4$ for $C(k,n)$ when $k=2,3$. Using $q$-series methods we derive closed formulas

  25. Wenyuan Zhang, Yu-Shen Liu, Zhizhong Han

    It is vital to infer a signed distance function (SDF) in multi-view based surface reconstruction. 3D Gaussian splatting (3DGS) provides a novel perspective for volume rendering, and shows advantages in rendering efficiency and quality. Although 3DGS provides a promising neural rendering option, it is still hard to infer SDFs for surface reconstruction with 3

  26. Nir Gadish

    We determine the 0-th Hochschild homology of the associative algebra of simplicial cochains valued in a PID: it consists of the ``finite-type" homotopy invariants of free loops, equivalently finite-type class functions on the fundamental group. One major motivation for this calculation is joint work in progress aiming to geometrically construct invariants of

  27. Lindsey M. Whitmore, Yalda Ramezani, Sumit Sharma, Michael R. Shirts

    The accurate treatment of long-range energy terms such as van der Waals interactions is crucial for reliable free energy calculations in molecular simulations. Methods like force switching, potential switching, potential shifting, and Ewald summation of van der Waals are commonly employed to smooth the truncation or otherwise manage these interactions at and

  28. Amy E. Miller, Zachary Slepian, Elizabeth A. Lada, Richard de Grijs

    We present a novel method for automatically detecting and characterising semi-resolved star clusters: clusters where the observational point-spread function (PSF) is smaller than the cluster's radius, but larger than the separations between individual stars. We apply our method to a 1.77 deg$^2$ field located in the Large Magellanic Cloud (LMC) using the VIS

  29. Felix Krones, Ben Walker, Terry Lyons, Adam Mahdi

    This work presents our team's (SignalSavants) winning contribution to the 2024 George B. Moody PhysioNet Challenge. The Challenge had two goals: reconstruct ECG signals from printouts and classify them for cardiac diseases. Our focus was the first task. Despite many ECGs being digitally recorded today, paper ECGs remain common throughout the world. Digitisin

  30. Mozhi Zhang, Pengyu Wang, Chenkun Tan, Mianqiu Huang

    Large Language Models (LLMs) acquire extensive knowledge and remarkable abilities from extensive text corpora, making them powerful tools for various applications. To make LLMs more usable, aligning them with human preferences is essential. Existing alignment techniques, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimiza

  31. Pengfei He, Zitao Li, Yue Xing, Yaling Li

    Zero-shot reasoning methods with Large Language Models (LLMs) offer significant advantages including great generalization to novel tasks and reduced dependency on human-crafted examples. However, the current zero-shot methods still have limitations in complex tasks, e.g., answering questions that require multi-step reasoning. In this paper, we address this l

  32. Yanhao Jin, Krishnakumar Balasubramanian, Lifeng Lai

    We investigate the in-context learning capabilities of transformers for the $d$-dimensional mixture of linear regression model, providing theoretical insights into their existence, generalization bounds, and training dynamics. Specifically, we prove that there exists a transformer capable of achieving a prediction error of order $\mathcal{O}(\sqrt{d/n})$ wit

  33. Juan Jose, Rojas-Constain

    I conducted a preregistered survey experiment (n=525) to assess the effectiveness of "accuracy nudges" against deepfakes (osf.io/69x17). The results, based on a sample of Colombian participants, replicated previous findings showing that prompting participants to assess the accuracy of a headline at the beginning of the survey significantly decreased their in

  34. Hansa Meghwani

    Ranking consistently emerges as a primary focus in information retrieval research. Retrieval and ranking models serve as the foundation for numerous applications, including web search, open domain QA, enterprise domain QA, and text-based recommender systems. Typically, these models undergo training on triplets consisting of binary relevance assignments, comp

  35. Yujun Zhou, Jingdong Yang, Yue Huang, Kehan Guo

    Artificial Intelligence (AI) is revolutionizing scientific research, yet its growing integration into laboratory environments presents critical safety challenges. Large language models (LLMs) and vision language models (VLMs) now assist in experiment design and procedural guidance, yet their "illusion of understanding" may lead researchers to overtrust unsaf

  36. Abeer Khan, Maria Hunaid Samiwala, Abeeha Zawar, Muhammad Qasim Pasta

    Cricket, a popular bat-and-ball game in South Asia, is played between two 11-player teams. The Pakistan Super League (PSL) is a commercial T20 domestic league comprised of six franchise-owned teams, where player selection is competitive. In this study, an existing role-based ranking structure is assessed that evaluates player performance in the context of te

  37. Taha Aksu, Chenghao Liu, Amrita Saha, Sarah Tan

    Time series forecasting aids decision-making, especially for stakeholders who rely on accurate predictions, making it very important to understand and explain these models to ensure informed decisions. Traditional explainable AI (XAI) methods, which underline feature or temporal importance, often require expert knowledge. In contrast, natural language explan

  38. Zifeng Zhu, Mengzhao Jia, Zhihan Zhang, Lang Li

    Multimodal Large Language Models (MLLMs) have demonstrated impressive abilities across various tasks, including visual question answering and chart comprehension, yet existing benchmarks for chart-related tasks fall short in capturing the complexity of real-world multi-chart scenarios. Current benchmarks primarily focus on single-chart tasks, neglecting the

  39. Younggeol Cho, Youngrae Kim, Junho Yoon, Seunghoon Hong

    Test-time adaptation (TTA) allows a model to be adapted to an unseen domain without accessing the source data. Due to the nature of practical environments, TTA has a limited amount of data for adaptation. Recent TTA methods further restrict this by filtering input data for reliability, making the effective data size even smaller and limiting adaptation poten

  40. Varun Murali, Guy Rosman, Sertac Karaman, Daniela Rus

    In this work, we consider the problem of learning end to end perception to control for ground vehicles solely from aerial imagery. Photogrammetric simulators allow the synthesis of novel views through the transformation of pre-generated assets into novel views.However, they have a large setup cost, require careful collection of data and often human effort to

  41. Andrea Dziubek, Kaibo Hu, Michael Karow, Michael Neunteufel

    We propose two parameter-robust mixed finite element methods for linear Cosserat elasticity. The Cosserat coupling constant $\mu_c$, connecting the displacement $u$ and rotation vector $\omega$, leads to possible locking phenomena in finite element methods. The formal limit of $\mu_c\to\infty$ enforces the constraint $\frac{1}{2}\operatorname{curl} u = \omeg

  42. Juan B. Pérez-Sánchez, Arghadip Koner, Sricharan Raghavan-Chitra, Joel Yuen-Zhou

    Molecular polaritons arise when the collective coupling between an ensemble of $N$ molecules and an optical mode exceeds individual photon and molecular linewidths. The complexity of their description stems from their multiscale nature, where the local dynamics on each molecule can, in principle, be influenced by the collective behavior of the entire ensembl

  43. Quang Dang, Murat Kucukosmanoglu, Michael Anoruo, Golshan Kargosha

    Assessing cognitive workload is crucial for human performance as it affects information processing, decision making, and task execution. Pupil size is a valuable indicator of cognitive workload, reflecting changes in attention and arousal governed by the autonomic nervous system. Cognitive events are closely linked to cognitive workload as they activate ment

  44. Daniel Liebau

    How will Decentralized Finance transform financial services? Using New Institutional Economics and Dynamic Capabilities Theory, I analyse survey data from 109 experts using non-parametric methods. Experts span traditional finance, DeFi industry, and academia. Four insights emerge: adoption expectations rise from negligible to 43% expecting at least high adop

  45. Fernando Saliby

    This study investigates the motion of a falling smartphone under the influence of air drag using acceleration data collected by its built-in accelerometer. The proper acceleration profiles demonstrate the suitability of the turbulent drag model in capturing the motion dynamics during both upward and downward phases. This approach provides an effective and ac

  46. Jiahao Qiu, Yifu Lu, Yifan Zeng, Jiacheng Guo

    Inference-time alignment enhances the performance of large language models without requiring additional training or fine-tuning but presents challenges due to balancing computational efficiency with high-quality output. Best-of-N (BoN) sampling, as a simple yet powerful approach, generates multiple responses and selects the best one, achieving improved perfo

  47. Kushagra Pandey, Jaideep Pathak, Yilun Xu, Stephan Mandt

    Diffusion models achieve state-of-the-art generation quality across many applications, but their ability to capture rare or extreme events in heavy-tailed distributions remains unclear. In this work, we show that traditional diffusion and flow-matching models with standard Gaussian priors fail to capture heavy-tailed behavior. We address this by repurposing

  48. Yiyan Xu, Wenjie Wang, Yang Zhang, Biao Tang

    Personalized content filtering, such as recommender systems, has become a critical infrastructure to alleviate information overload. However, these systems merely filter existing content and are constrained by its limited diversity, making it difficult to meet users' varied content needs. To address this limitation, personalized content generation has emerge

  49. Ho Sung Shim, Hyoungjun Park, Kyuhan Lee, Jang-Sun Park

    Smishing, which aims to illicitly obtain personal information from unsuspecting victims, holds significance due to its negative impacts on our society. In prior studies, as a tool to counteract smishing, machine learning (ML) has been widely adopted, which filters and blocks smishing messages before they reach potential victims. However, a number of challeng

  50. Ange Lou, Benjamin Planche, Zhongpai Gao, Yamin Li

    Numerous recent approaches to modeling and re-rendering dynamic scenes leverage plane-based explicit representations, addressing slow training times associated with models like neural radiance fields (NeRF) and Gaussian splatting (GS). However, merely decomposing 4D dynamic scenes into multiple 2D plane-based representations is insufficient for high-fidelity

  51. Travis Cuvelier, Sean Ha, Maretta Morovitz

    We consider the case where an adversary is conducting a surveillance campaign against a networked control system (NCS), and take the perspective of a defender/control system operator who has successfully isolated the cyber intruder. To better understand the adversary's intentions and to drive up their operating costs, the defender directs the adversary towar

  52. Jiajing Chen, Runyuan Bao, Hongye Zheng, Zhen Qi

    This study aims to improve the accuracy and quality of large-scale language models (LLMs) in answering questions by integrating Elasticsearch into the Retrieval Augmented Generation (RAG) framework. The experiment uses the Stanford Question Answering Dataset (SQuAD) version 2.0 as the test dataset and compares the performance of different retrieval methods,

  53. Nan Xu, Xuezhe Ma

    Interestingly, LLMs yet struggle with some basic tasks that humans find trivial to handle, e.g., counting the number of character r's in the word "strawberry". There are several popular conjectures (e.g., tokenization, architecture and training data) regarding the reason for deficiency of LLMs in simple word-based counting problems, sharing the similar belie

  54. Chihang Wang, Yuxin Dong, Zhenhong Zhang, Ruotong Wang

    This paper focuses on the development of an advanced intelligent article scoring system that not only assesses the overall quality of written work but also offers detailed feature-based scoring tailored to various article genres. By integrating the pre-trained BERT model with the large language model Chat-GPT, the system gains a deep understanding of both th

  55. Sébastien Henry, John A. Christian

    We propose a modified normalized direct linear transform (DLT) algorithm for solving the perspective-n-point (PnP) problem with much better behavior than the conventional DLT. The modification consists of analytically weighting the different measurements in the linear system with a negligible increase in computational load. Our approach exhibits clear improv

  56. Rhui Dih Lee, Laura Wynter

    We address the question of how to successively add new knowledge to an LLM whilst retaining previously-added knowledge. We consider two settings, semi-cooperative and fully-cooperative. Overall, LoRA performs better in most cases than full-fine tuning of all parameters when both new knowledge acquisition and retention of old, including recent, knowledge are

  57. Santanu S Dey, Dahye Han, Yang Wang

    In this paper, we study the strength of convex relaxations obtained by convexification of aggregation of constraints for a set $S$ described by two bilinear bipartite equalities. Aggregation is the process of rescaling the original constraints by scalar weights and adding the scaled constraints together. It is natural to study the aggregation technique as it

  58. Sergei P. Maydanyuk, Gyorgy Wolf

    We investigate production of electron-positron pairs (dileptons) in the scattering of protons off nuclei. Focus is directed on clarifying role of nuclear interactions which make basis in mechanisms of scattering process and structure of nuclei. For that, we constructed a new model of production of dileptons, where scattering of nuclei and their structure are

  59. Renguang Chen, Guolong Zheng, Xu Yang, Zhide Chen

    The growing popularity of online sports and exercise necessitates effective methods for evaluating the quality of online exercise executions. Previous action quality assessment methods, which relied on labeled scores from motion videos, exhibited slightly lower accuracy and discriminability. This limitation hindered their rapid application to newly added exe

  60. Nathan J. Di Vaira, Lukasz Laniewski-Wollk, Raymond L. Johnson, Saiied M. Aminossadati

    This work is the first computational study of proppant leak-off through coal cleats that accounts for proppant retention in cleats, occlusion formation at cleat entrances, the resulting control of fluid leak-off, and the influence of realistic cleat roughness on these factors. Suspensions are simulated with a coupled lattice Boltzmann method-discrete element

  61. Héctor Laria, Alex Gomez-Villa, Kai Wang, Bogdan Raducanu

    Recent advances in diffusion models have significantly enhanced image generation capabilities. However, customizing these models with new classes often leads to unintended consequences that compromise their reliability. We introduce the concept of open-world forgetting to characterize the vast scope of these unintended alterations. Our work presents the firs

  62. Shuyang Wang, Diego Klabjan

    Recent work by Woodworth et al. (2020) shows that the optimization dynamics of gradient descent for overparameterized problems can be viewed as low-dimensional dual dynamics induced by a mirror map, explaining the implicit regularization phenomenon from the mirror descent perspective. However, the methodology does not apply to algorithms where update directi

  63. Jiacheng Ye, Jiahui Gao, Shansan Gong, Lin Zheng

    Autoregressive language models, despite their impressive capabilities, struggle with complex reasoning and long-term planning tasks. We introduce discrete diffusion models as a novel solution to these challenges. Through the lens of subgoal imbalance, we demonstrate how diffusion models effectively learn difficult subgoals that elude autoregressive approache

  64. M. A. Rayhan, M. M. Hossain, M. M. Uddin, S. H. Naqib

    Double perovskite halides are promising materials for renewable energy production, meeting the criteria to address energy scarcity issues. As a result, studying these halides could be useful for optoelectronic and solar cell applications. In this study, we investigated the structural, mechanical, thermodynamic, electronic, and optical properties of A2AgIrCl6

  65. Wei Jie Yeo, Ranjan Satapathy, Erik Cambria

    Large Language Models (LLMs) are capable of generating persuasive Natural Language Explanations (NLEs) to justify their answers. However, the faithfulness of these explanations should not be readily trusted at face value. Recent studies have proposed various methods to measure the faithfulness of NLEs, typically by inserting perturbations at the explanation

  66. Muhe Ding, Yang Ma, Pengda Qin, Jianlong Wu

    Multimodal Large Language Models (MLLMs) have recently received substantial interest, which shows their emerging potential as general-purpose models for various vision-language tasks. MLLMs involve significant external knowledge within their parameters; however, it is challenging to continually update these models with the latest knowledge, which involves hu

  67. Gaoyang Pang, Wanchun Liu, Dusit Niyato, Daniel Quevedo

    Wireless Human-Machine Collaboration (WHMC) represents a critical advancement for Industry 5.0, enabling seamless interaction between humans and machines across geographically distributed systems. As the WHMC systems become increasingly important for achieving complex collaborative control tasks, ensuring their stability is essential for practical deployment

  68. Jiarui Ji, Yang Li, Hongtao Liu, Zhicheng Du

    Public scarce resource allocation plays a crucial role in economics as it directly influences the efficiency and equity in society. Traditional studies including theoretical model-based, empirical study-based and simulation-based methods encounter limitations due to the idealized assumption of complete information and individual rationality, as well as const

  69. Yangyang Zhang, Zhenwei Li, Xuefei Chen, Zhanwen Han

    Double helium white dwarfs (He WDs) are one type of gravitational wave source and are greatly important in the studies of binary interaction, particularly in the common envelope (CE) ejection physics. Most double He WDs with mass ratios of q~1 are formed through a particular channel. In this channel, one He WD is initially produced from a red giant (RG) with

  70. Xiaoyong Huang, Heli Sun, Qunshu Gao, Wenjie Huang

    With the rapid development of the internet, the richness of User-Generated Contentcontinues to increase, making Multimodal Aspect-Based Sentiment Analysis (MABSA) a research hotspot. Existing studies have achieved certain results in MABSA, but they have not effectively addressed the analytical challenges in scenarios where multiple entities and sentiments co

  71. Russelle Guadalupe

    In this paper, we study the $5$-dissections of certain Ramanujan's theta functions, particularly $\psi(q)\psi(q^2), \varphi(-q)$ and $\varphi(-q)\varphi(-q^2)$, and derive an identity for $q(q;q)_{\infty}^6/(q^5;q^5)_{\infty}^6$ in terms of certain products of the Rogers-Ramanujan continued fraction $R(q)$. Using this identity, we give another proof of the m

  72. Chenhang Cui, An Zhang, Yiyang Zhou, Zhaorun Chen

    The recent advancements in large language models (LLMs) and pre-trained vision models have accelerated the development of vision-language large models (VLLMs), enhancing the interaction between visual and linguistic modalities. Despite their notable success across various domains, VLLMs face challenges in modality alignment, which can lead to issues like hal

  73. Grace Luo, Christopher Boyer, Siddharth Penmetsa

    In this paper, we study a theoretical math problem of game theory and calculus of variations in which we minimize a functional involving two players. A general relationship between the optimal strategies for both players is presented, followed by computer analysis as well as polynomial approximation. Nash equilibrium strategies are determined through algebra

  74. Jiahao Wang, Amer Shalaby

    Public transit systems play a crucial role in providing efficient and sustainable transportation options in urban areas. However, these systems face various challenges in meeting commuters' needs. On the other hand, despite the rapid development of Large Language Models (LLMs) worldwide, their integration into transit systems remains relatively unexplored. T

  75. Yanming Zhang, Akshith Kota, Eric Papenhausen, Klaus Mueller

    Causal networks are widely used in many fields to model the complex relationships between variables. A recent approach has sought to construct causal networks by leveraging the wisdom of crowds through the collective participation of humans. While this can yield detailed causal networks that model the underlying phenomena quite well, it requires a large numb

  76. June M. Liu, He Cao, Renliang Sun, Rui Wang

    Generating emotionally appropriate responses in conversations with large language models presents a significant challenge due to the complexities of human emotions and cognitive processes, which remain largely underexplored in their critical role in social interactions. In this study, we introduce a two-stage automatic data generation framework to create CAP

  77. Chenyang Zhang, Jiayi Lin, Haibo Tong, Bingxuan Hou

    Large language models (LLMs) show remarkable abilities with instruction tuning. However, they fail to achieve ideal tasks when lacking high-quality instruction tuning data on target tasks. Multi-Aspect Controllable Text Generation (MCTG) is a representative task for this dilemma, where aspect datasets are usually biased and correlated. Existing work exploits

  78. Muhe Ding, Jianlong Wu, Xue Dong, Xiaojie Li

    Knowledge distillation is a mainstream algorithm in model compression by transferring knowledge from the larger model (teacher) to the smaller model (student) to improve the performance of student. Despite many efforts, existing methods mainly investigate the consistency between instance-level feature representation or prediction, which neglects the category

  79. Tianqing Zhou, Bobo Wang, Dong Qin, Xuefang Nie

    Cache-assisted ultra-dense mobile edge computing (MEC) networks are a promising solution for meeting the increasing demands of numerous Internet-of-Things mobile devices (IMDs). To address the complex interferences caused by small base stations (SBSs) deployed densely in such networks, this paper explores the combination of orthogonal frequency division mult

  80. K. Mahesh Krishna

    Motivated from Deutsch entropic uncertainty principle and several product uncertainty principles, we derive an uncertainty principle for the product of entropies using functions.

  81. Sabit Hassan, Hye-Young Chung, Xiang Zhi Tan, Malihe Alikhani

    When assisting people in daily tasks, robots need to accurately interpret visual cues and respond effectively in diverse safety-critical situations, such as sharp objects on the floor. In this context, we present M-CoDAL, a multimodal-dialogue system specifically designed for embodied agents to better understand and communicate in safety-critical situations.

  82. Alyssa Cassity, Hieu Le, Hernan Santos, Erik Priest

    The study focuses on developing a digital twin testbed tailored for public safety technologies, incorporating simulated wireless communication within the digital world. The integration enables rapid analysis of signal strength, facilitating effective communication among personnel during catastrophic incidents in the virtual environment. The virtual world als

  83. M. Hasan Barbhuiya, Paul A. Cassak, Alex Chasapis, Michael A. Shay

    Magnetic reconnection often initiates abruptly and then rapidly progresses to a nonlinear quasi-steady state. While satellites frequently detect reconnection events, ascertaining whether the system has achieved steady-state or is still evolving in time remains challenging. Here, we propose that the relatively rapid opening of reconnection separatrices within

  84. Jingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu

    Large vision-language models (LVLMs) have witnessed significant progress on visual understanding tasks. However, they often prioritize language knowledge over image information on visual reasoning tasks, incurring performance degradation. To tackle this issue, we first identify the drawbacks of existing solutions (i.e., limited multi-modal reasoning capaciti

  85. Diyi Liu, Ankur Shiledar, Hyeonsup Lim, Vivek Sujan

    Understanding the dynamics of truck volumes and activities across the skeleton traffic network is pivotal for effective traffic planning, traffic management, sustainability analysis, and policy making. Yet, relying solely on average annual daily traffic volume for trucks cannot capture the temporal changes over time. Recently, the Traffic Monitoring Analysis

  86. Shaoming Xu, Arvind Renganathan, Ankush Khandelwal, Rahul Ghosh

    Streamflow, vital for water resource management, is governed by complex hydrological systems involving intermediate processes driven by meteorological forces. While deep learning models have achieved state-of-the-art results of streamflow prediction, their end-to-end single-task learning approach often fails to capture the causal relationships within these s

  87. Yuan Li, Zicheng Ye, Huazi Zhang, Jun Wang

    In this paper, we study the system-level advantages provided by rateless coding, early termination and power allocation strategy for multiple users distributed across multiple cells. In a multi-cell scenario, the early termination of coded transmission not only reduces finite-length loss akin to the single-user scenario but also yields capacity enhancements

  88. Kavinayan P. Sivakumar, Yi Shen, Zachary Bell, Scott Nivison

    In this paper, we study an inverse reinforcement learning problem that involves learning the reward function of a learning agent using trajectory data collected while this agent is learning its optimal policy. To address this problem, we propose an inverse reinforcement learning method that allows us to estimate the policy parameters of the learning agent wh

  89. Sidi Wu

    Physics-Informed Neural Networks (PINNs) have emerged as powerful tools for solving partial differential equations (PDEs). However, training PINNs from scratch is often computationally intensive and time-consuming. To address this problem, we propose a parameter-efficient approach that fine-tunes pre-trained DeepONet models within the PINN framework (FTO-PIN

  90. Likun Xie

    We prove lower bounds for the number of primes $p \leq N + b$ such that $p-b$ is divisible by $2^{k(N)}$ and has at most $k$ odd prime factors ($k \geq 2$), assuming $2^{k(N)} \leq N^\theta$ for some $\theta > 0$ depending on $k$. The proof uses a variant of Chen's method, weighted sieves, and Elliott's results on primes in arithmetic progressions with large

  91. Nghia Hieu Nguyen, Tho Thanh Quan, Ngan Luu-Thuy Nguyen

    Text-based VQA is a challenging task that requires machines to use scene texts in given images to yield the most appropriate answer for the given question. The main challenge of text-based VQA is exploiting the meaning and information from scene texts. Recent studies tackled this challenge by considering the spatial information of scene texts in images via e

  92. Aimina Ali Eli, Abida Ali

    Medical image analysis has emerged as an essential element of contemporary healthcare, facilitating physicians in achieving expedited and precise diagnosis. Recent breakthroughs in deep learning, a subset of artificial intelligence, have markedly revolutionized the analysis of medical pictures, improving the accuracy and efficiency of clinical procedures. De

  93. Yulun Xu, Kai Zheng

    Let $D$ be a smooth divisor on a closed K\"ahler manifold $X$. First, we prove that Poincar\'e type constant scalar curvature K\"ahler (cscK) metric with a singularity at $D$ is unique up to a holomorphic transformation on $X$ that preserves $D$, if there are no nontrivial holomorphic vector fields on $D$. For the general case, we propose a conjecture relati

  94. Zongbin Chen

    We give a proof of the geometric fundamental lemma of Kottwitz. As explained by Laumon, this implies the fundamental lemma for the unitary groups.

  95. Russel Arbore, Jeffrey Liu, Aidan Wefel, Steven Gao

    Voxels are a geometric representation used for rendering volumes, multi-resolution models, and indirect lighting effects. Since the memory consumption of uncompressed voxel volumes scales cubically with resolution, past works have introduced data structures for exploiting spatial sparsity and homogeneity to compress volumes and accelerate ray tracing. Howeve

  96. Xiping Liu, Zhao Tan

    Text-to-SQL parsing involves the translation of natural language queries (NLQs) into their corresponding SQL commands. A principal challenge within this domain is the formulation of SQL queries that are not only syntactically correct but also semantically aligned with the natural language input. However, the intrinsic disparity between the NLQ and the SQL po

  97. Eli N. Weinstein, Elizabeth B. Wood, David M. Blei

    A central question in human immunology is how a patient's repertoire of T cells impacts disease. Here, we introduce a method to infer the causal effects of T cell receptor (TCR) sequences on patient outcomes using observational TCR repertoire sequencing data and clinical outcomes data. Our approach corrects for unobserved confounders, such as a patient's env

  98. Dandan Chen, Ziyin Zou

    Recently Andrews and Bachraoui proved identities relating certain restricted partitions into distinct even parts with restricted 4-regular partitions by the theory of basic hypergeometric series. They also posed a question regarding combinatorial proofs for these results. In this paper, we establish bijections to provide combinatorial proofs for these result

  99. Nirmali Roy, Anuradha Jha

    In this article, we study a two-dimensional singularly perturbed parabolic equation of the convection-diffusion type, characterized by discontinuities in the source term and convection coefficient at a specific point in the domain. These discontinuities lead to the development of interior layers. To address these layers and ensure uniform convergence, we pro

  100. Noah Toyonaga, L Mahadevan

    We introduce an additive approach for the design of a class of transformable structures based on two-bar linkages ("scissor mechanisms") joined at vertices to form a two dimensional lattice. Our discussion traces an underlying mathematical similarity between linkage mechanisms, origami, and kirigami and inspires our name for these structures: karigami. We sh