Skip to content

May 2025 arXiv papers — page 62

Showing 6,1016,200 of 24,552 papers

  1. Buse Sibel Korkmaz, Rahul Nair, Elizabeth M. Daly, Antonio del Rio Chanona

    Current debiasing approaches often result a degradation in model capabilities such as factual accuracy and knowledge retention. Through systematic evaluation across multiple benchmarks, we demonstrate that existing debiasing methods face fundamental trade-offs, particularly in smaller models, leading to reduced truthfulness, knowledge loss, or unintelligible

  2. Víctor Arnaiz

    In this article we study the semiclassical asymptotics of the Martinet sub-Laplacian on the flat toroidal cylinder $M = \mathbb{R} \times \mathbb{T}^2$. We describe the asymptotic distribution of sequences of eigenfunctions oscillating at different scales prefixed by Rothschild-Stein estimates via the introduction of adapted two-microlocal semiclassical meas

  3. Bhanuka Gamage, Thanh-Toan Do, Nicholas Seow Chiang Price, Arthur Lowery

    Over the last decade there has been considerable research into how artificial intelligence (AI), specifically computer vision, can assist people who are blind or have low-vision (BLV) to understand their environment. However, there has been almost no research into whether the tasks (object detection, image captioning, text recognition etc.) and devices (smar

  4. Luca Sandrock, Thomas Schick

    A cohomology class u of a topological space X is atoroidal if its pullback to the torus vanishes for every map from a torus to X. Furthermore, X is atoroidally symplectic if there is an atoroidal cohomology class $u\in H^2(X;F)$ such that $u^n$ is non-zero. We prove that every atoroidally symplectic CW-complex X of dimension 2n has topological complexity 4n.

  5. Ahmad M. Nazar, Mohamed Y. Selim, Daji Qiao, Hongwei Zhang

    Artificial intelligence (AI) and wireless networking advancements have created new opportunities to enhance network efficiency and performance. In this paper, we introduce Next-Generation GPT (NextG-GPT), an innovative framework that integrates retrieval-augmented generation (RAG) and large language models (LLMs) within the wireless systems' domain. By lever

  6. P. A. Solar, B. Reinoso, D. R. G. Schleicher, R. S. Klessen

    The origin of supermassive black holes is an open question that has been explored considering gas- and collision-based formation channels to explain the high number of quasars observed in the early Universe. According to numerical simulations, supermassive stars can be formed in atomic cooling halos when protostars reach accretion rates greater than $\sim 10

  7. Michail Spitieris, Massimiliano Ruocco, Abdulmajid Murad, Alessandro Nocente

    Recent advances in generative AI offer promising solutions for synthetic data generation but often rely on large datasets for effective training. To address this limitation, we propose a novel generative model that learns from limited data by incorporating physical constraints to enhance performance. Specifically, we extend the VAE architecture by incorporat

  8. Qiang Hu, Qimei Wang, Jia Chen, Xuantao Ji

    White Light Imaging (WLI) and Narrow Band Imaging (NBI) are the two main colonoscopic modalities for polyp classification. While NBI, as optical chromoendoscopy, offers valuable vascular details, WLI remains the most common and often the only available modality in resource-limited settings. However, WLI-based methods typically underperform, limiting their cl

  9. Subham Dutta, Ahmed Atteya, Pralay Kumar Karmakar

    We present a comprehensive overview of the formation mechanism of plasma fireball sheath (PFS) structures, sheath-induced collective phenomena, associated relevant instabilities, and corresponding onset conditions. It includes an optimum set of self-illustrative schematic figures, relevantly manifesting the instability triggering dynamics, various involved m

  10. Tin Trung Nguyen, Jiannan Xu, Zora Che, Phuong-Anh Nguyen-Le

    Although popularized AI fairness metrics, e.g., demographic parity, have uncovered bias in AI-assisted decision-making outcomes, they do not consider how much effort one has spent to get to where one is today in the input feature space. However, the notion of effort is important in how Philosophy and humans understand fairness. We propose a philosophy-inform

  11. Rex Chen, Stephanie Milani, Zhicheng Zhang, Norman Sadeh

    Poor interpretability hinders the practical applicability of multi-agent reinforcement learning (MARL) policies. Deploying interpretable surrogates of uninterpretable policies enhances the safety and verifiability of MARL for real-world applications. However, if these surrogates are to interact directly with the environment within human supervisory framework

  12. Farid Najar, Dominique Barth, Yann Strozecki

    Combinatorial optimization (CO) problems are traditionally addressed using Operations Research (OR) methods, including metaheuristics. In this study, we introduce a demand selection problem for the Vehicle Routing Problem (VRP) with an emission quota, referred to as QVRP. The objective is to minimize the number of omitted deliveries while respecting the poll

  13. Helin Wang, Jiarui Hai, Dongchao Yang, Chen Chen

    Target Speech Extraction (TSE) aims to isolate a target speaker's voice from a mixture of multiple speakers by leveraging speaker-specific cues, typically provided as auxiliary audio (a.k.a. cue audio). Although recent advancements in TSE have primarily employed discriminative models that offer high perceptual quality, these models often introduce unwanted a

  14. Marta Aparicio Rodriguez, Xenia Miscouridou, Anastasia Borovykh

    Despite significant advances in quality and complexity of the generations in text-to-image models, prompting does not always lead to the desired outputs. Controlling model behaviour by directly steering intermediate model activations has emerged as a viable alternative allowing to reach concepts in latent space that may otherwise remain inaccessible by promp

  15. Zirui Li, Siwei Wu, Yizhi Li, Xingyu Wang

    The rapid advancement of unsupervised representation learning and large-scale pre-trained vision-language models has significantly improved cross-modal retrieval tasks. However, existing multi-modal information retrieval (MMIR) studies lack a comprehensive exploration of document-level retrieval and suffer from the absence of cross-domain datasets at this gr

  16. Maharshi A. Sharma, Albert E. Patterson

    Improved additive manufacturing capabilities are vital for the future development and improvement of ubiquitous robotic systems. These machines can be integrated into existing robotic systems to allow manufacturing and repair of components, as well as fabrication of custom parts for the robots themselves. The fused filament fabrication (FFF) process is one o

  17. Robin D. Pesl, Jerin G. Mathew, Massimo Mecella, Marco Aiello

    Integrating multiple (sub-)systems is essential to create advanced Information Systems. Difficulties mainly arise when integrating dynamic environments, e.g., the integration at design time of not yet existing services. This has been traditionally addressed using a registry that provides the API documentation of the endpoints. Large Language Models have show

  18. Yun Zhao, Harry Zheng

    Mixed optimal stopping and stochastic control problems define variational inequalities with non-linear Hamilton-Jacobi-Bellman (HJB) operators, whose numerical solution is notoriously difficult and lack of reliable benchmarks. We first use the dual approach to transform it into a linear operator, and then introduce a Fractional-Boundary-Regularized Deep Gale

  19. J. U. Ness, N. Schartel, M. Santos-Lleo

    Novel studies are presented demonstrating that the data of ESA's XMM-Newton mission are efficiently used by an engaged and productive community. 87% of the available time budget during the reference period 2000-2024 of 556Ms was used in at least one of 8486 publications (84% of 16894 observations) with a re-use of a factor up to 15 in dedicated publications.

  20. João Coelho, Bruno Martins, João Magalhães, Chenyan Xiong

    Neural retrieval models excel in Web search, but their training requires substantial amounts of labeled query-document pairs, which are costly to obtain. With the widespread availability of Web document collections like ClueWeb22, synthetic queries generated by large language models offer a scalable alternative. Still, synthetic training queries often vary i

  21. Weiming Zhi, Ziyong Ma, Tianyi Zhang, Matthew Johnson-Roberson

    Autonomous robots typically need to construct representations of their surroundings and adapt their motions to the geometry of their environment. Here, we tackle the problem of constructing a policy model for collision-free motion generation, consistent with the environment, from a single input RGB image. Extracting 3D structures from a single image often in

  22. Roman Beltiukov, Karthik Bhattaram, Evania Cheng, Vinod Kanigicherla

    With the worldwide growth of remote communication and telepresence, network measurements form a cornerstone of effective performance assessment and diagnostics for Internet users. Most often, users seek for overall connection performance measurement using publicly available tools (also known as `speed tests') that provide an overview of their connection's th

  23. Maximilian Heisinger, Clemens Hofstadler

    We present f4ncgb, a new open-source C++ library for Gr\"obner basis computations in free algebras, which transfers recent advancements in commutative Gr\"obner basis software to the noncommutative setting. As our experiments show, f4ncgb establishes a new state-of-the-art for noncommutative Gr\"obner basis computations. We also discuss implementation detail

  24. Victor Bailey, Deguang Han, Keri Kornelson, David Larson

    The theory of dynamical frames evolved from practical problems in dynamical sampling where the initial state of a vector needs to be recovered from the space-time samples of evolutions of the vector. This leads to the investigation of structured frames obtained from the orbits of evolution operators. One of the basic problems in dynamical frame theory is to

  25. Kapil Vaidya, Abishek Sankararaman, Jialin Ding, Chuan Lei

    NL2SQL (natural language to SQL) systems translate natural language into SQL queries, allowing users with no technical background to interact with databases and create tools like reports or visualizations. While recent advancements in large language models (LLMs) have significantly improved NL2SQL accuracy, schema ambiguity remains a major challenge in enter

  26. Ken Huang, Vineeth Sai Narajala, John Yeoh, Jason Ross

    Traditional Identity and Access Management (IAM) systems, primarily designed for human users or static machine identities via protocols such as OAuth, OpenID Connect (OIDC), and SAML, prove fundamentally inadequate for the dynamic, interdependent, and often ephemeral nature of AI agents operating at scale within Multi Agent Systems (MAS), a computational sys

  27. Junnan Liu, Linhao Luo, Thuy-Trang Vu, Gholamreza Haffari

    Recent advances in large language models (LLMs) demonstrate their impressive reasoning capabilities. However, the reasoning confined to internal parametric space limits LLMs' access to real-time information and understanding of the physical world. To overcome this constraint, we introduce SituatedThinker, a novel framework that enables LLMs to ground their r

  28. Lingjun Zhao, Hal Daumé

    Faithful free-text explanations are important to ensure transparency in high-stakes AI decision-making contexts, but they are challenging to generate by language models and assess by humans. In this paper, we present a measure for Prediction-EXplanation (PEX) consistency, by extending the concept of weight of evidence. This measure quantifies how much a free

  29. Theo Diamandis, Tarun Chitra, Guillermo Angeris

    As markets have digitized, the number of tradable products has skyrocketed. Algorithmically constructed portfolios of these assets now dominate public and private markets, resulting in a combinatorial explosion of tradable assets. In this paper, we provide a simple means to compute market clearing prices for semi-fungible assets which have a partial ordering

  30. Valerii Startsev, Alexander Ustyuzhanin, Alexey Kirillov, Dmitry Baranchuk

    Pre-training equips text-to-image (T2I) models with broad world knowledge, but this alone is often insufficient to achieve high aesthetic quality and alignment. Consequently, supervised fine-tuning (SFT) is crucial for further refinement. However, its effectiveness highly depends on the quality of the fine-tuning dataset. Existing public SFT datasets frequen

  31. Mohammad Jafari

    We investigate a network model in which a single random walker combines local diffusion with preferential resetting to previously visited nodes. Each arrival deposits one unit of stress on the target node, and threshold crossings trigger sandpile-like relaxation cascades. The fixed per-neighbor transfer rule produces a brittle transition on Watts--Strogatz n

  32. Adriano De Santana, Rene Baltazar, Robson Vinciguerra, Wilian De Araujo

    This paper investigates the isotropy groups of derivations on the Quantum Plane $\Bbbk_q[x, y]$, defined by the relation $yx = qxy$, where $q \in \Bbbk^*$, with $q^2\neq 1$. The main goal is to determine the automorphisms of the Quantum Plane that commutes with a fixed derivation $\delta$. We describe conditions under which the isotropy group $\text{Aut}_\de

  33. Ziyang Ma, Xiquan Li, Yakun Song, Wenxi Chen

    Recent advancements in large audio language models (LALMs) have demonstrated impressive results and promising prospects in universal understanding and reasoning across speech, music, and general sound. However, these models still lack the ability to recognize their knowledge boundaries and refuse to answer questions they don't know proactively. While there h

  34. Arianna Ceccarelli, Alexander P. Browning, Tai Chaiamarit, Ilan Davis

    Advances in experimental techniques allow the collection of high-resolution spatio-temporal data that track individual motile entities. These tracking data can be used to calibrate mathematical models describing the motility of individual entities. The challenges in calibrating models for single-agent motion derive from the intrinsic characteristics of exper

  35. Kazi Mahathir Rahman, Showrin Rahman, Sharmin Sultana Srishty

    Text-embedded image generation plays a critical role in industries such as graphic design, advertising, and digital content creation. Text-to-Image generation methods leveraging diffusion models, such as TextDiffuser-2, have demonstrated promising results in producing images with embedded text. TextDiffuser-2 effectively generates bounding box layouts that g

  36. Ahmadreza Montazerolghaem, Somaye Imanpour

    Software-defined networking (SDN) represents a revolutionary shift in network technology by decoupling the data plane from the control plane.}In this architecture, all network decision-making processes are centralized in a controller, meaning each switch receives routing information from the controller and forwards network packets accordingly. This clearly h

  37. H. Hajaiej, R. Leitao

    In this paper, we establish an anisotropic version of Campanato Theorem and show that the anisotropic Bessel spaces are continuously embedded in the spaces of Holder continuous functions. As an application of this embedding, we build fundamental solutions for a class of anisotropic fractional Laplacian operators.

  38. Jimeng Shi, Sizhe Zhou, Bowen Jin, Wei Hu

    Large language models (LLMs) often need to incorporate external knowledge to solve theme-specific problems. Retrieval-augmented generation (RAG) has shown its high promise, empowering LLMs to generate more qualified responses with retrieved external data and knowledge. However, most RAG methods retrieve relevant documents based on either sparse or dense retr

  39. Justice Akuoko-Frimpong, Edward Shao, Jonathan Ta

    Traditional regression models assume stationary relationships between predictors and responses, failing to capture the spatial heterogeneity present in many environmental, epidemiological, and ecological processes. To address this limitation, we develop a scalable Bayesian framework for spatially varying coefficient (SVC) models, implemented in the \pkg{svc}

  40. Utkarsh Sahu, Zhisheng Qi, Yongjia Lei, Ryan A. Rossi

    Large language models have been extensively studied as neural knowledge bases for their knowledge access, editability, reasoning, and explainability. However, few works focus on the structural patterns of their knowledge. Motivated by this gap, we investigate these structural patterns from a graph perspective. We quantify the knowledge of LLMs at both the tr

  41. Thomas Rüd

    To answer a question about the distribution of products of elliptic curves in isogeny classes of abelian surfaces defined over finite fields, we compute specific orbital integrals in the group $\mathrm{GSp}_4$. More precisely, we compute integrals over the orbits of elements in the subgroup $\mathrm{GL}_2\times_{\det} \mathrm{GL}_2$. As a first step towards

  42. Sahel Sharifymoghaddam, Ronak Pradeep, Andre Slavescu, Ryan Nguyen

    The adoption of large language models (LLMs) as rerankers in multi-stage retrieval systems has gained significant traction in academia and industry. These models refine a candidate list of retrieved documents, often through carefully designed prompts, and are typically used in applications built on retrieval-augmented generation (RAG). This paper introduces

  43. Zeinab Lashkaripour, Masoud Khosravi-Farmad, AhmadReza Montazerolghaem, Razieh Rezaee

    IoT is a dynamic network of interconnected things that communicate and exchange data, where security is a significant issue. Previous studies have mainly focused on attack classifications and open issues rather than presenting a comprehensive overview on the existing threats and vulnerabilities. This knowledge helps analyzing the network in the early stages

  44. Thor E. Andreassen, Taylor P. Trentadue, Andrew R. Thoreson, Kai-Nan An

    While computational modeling may help to develop new treatment options for hand and wrist injuries, at present, few models exist. The time and expertise required to develop and use these models is considerable. Moreover, most do not allow for variation of material properties, instead relying on literature reported averages. We have developed a novel automate

  45. Yuzheng Hu, Fan Wu, Haotian Ye, David Forsyth

    Online reinforcement learning (RL) excels in complex, safety-critical domains but suffers from sample inefficiency, training instability, and limited interpretability. Data attribution provides a principled way to trace model behavior back to training samples, yet existing methods assume fixed datasets, which is violated in online RL where each experience bo

  46. Vivek Pandey, Sudhir K. Pandey

    The anomalous Nernst conductivity (ANC) is a key transport property in magnetic and topological materials, arising from the Berry curvature ($\boldsymbol\Omega$) of electronic bands. It offers deep insight into the underlying topology and thermoelectric behavior. While Wannier interpolation have become popular for calculating ANC due to their computational e

  47. H. Choudhury, A. Battey, C. Paz-Soldan, J. Lestz

    Helicon waves satisfying the normal wave-particle cyclotron resonance are observed to limit the growth and maximum energy of relativistic electrons (REs) in low-density Ohmic DIII-D tokamak plasmas. Following the application of helicon waves, pitch-angle scattering of high-energy REs causes an increase in both synchrotron and electron-cyclotron emissions. Th

  48. Shixuan Zhang, Suhan Zhong

    We propose moment relaxations for data-driven Wasserstein distributionally robust optimization problems. Conditions are identified to ensure asymptotic consistency of such relaxations for both single-stage and two-stage problems, together with examples that illustrate their necessity. Numerical experiments are also included to illustrate the proposed relaxat

  49. Ibukun Olatunji, Mark Sheppard

    This paper argues that token prediction is fundamentally misaligned with real creativity. While next-token models have enabled impressive advances in language generation, their architecture favours surface-level coherence over spontaneity, originality, and improvisational risk. We use battle rap as a case study to expose the limitations of predictive systems

  50. Vasily Melnikov

    We introduce a new paradigm for risk sharing that generalizes earlier models based on discrete agents and extends them to allow for sharing risk within a continuum of agents. Agents are represented by points of a measure space and have potentially heterogeneous risk preferences modeled by risk measures on a separable probability space. We derive the dual rep

  51. James P. Crutchfield, Alexandra Jurgens

    We develop information theory for the temporal behavior of memoryful agents moving through complex -- structured, stochastic -- environments. We introduce and explore information processes -- stochastic processes produced by cognitive agents in real-time as they interact with and interpret incoming stimuli. We provide basic results on the ergodicity and sema

  52. Manveer Singh Tamber, Suleman Kazi, Vivek Sourabh, Jimmy Lin

    We investigate improving the retrieval effectiveness of embedding models through the lens of corpus-specific fine-tuning. Prior work has shown that fine-tuning with queries generated using a dataset's retrieval corpus can boost retrieval effectiveness for the dataset. However, we find that surprisingly, fine-tuning using the conventional InfoNCE contrastive

  53. Giuseppe Ruggiero, Matteo Testa, Jurgen Van de Walle, Luigi Di Caro

    Self-supervised learning (SSL) has reduced the reliance on expensive labeling in speech technologies by learning meaningful representations from unannotated data. Since most SSL-based downstream tasks prioritize content information in speech, ideal representations should disentangle content from unwanted variations like speaker characteristics in the SSL rep

  54. Maria Spethmann, Peter Stano, Daniel Loss

    Across most qubit platforms, the readout fidelities do not keep up with the gate fidelities, and new ways to increase the readout fidelities are searched for. For semiconductor spin qubits, a typical qubit-readout signal consists of a finite stretch of a digitized charge-sensor output. Such a signal trace is usually analyzed by compressing it into a single v

  55. Xun Deng, Sicheng Zhong, Barış Bayazıt, Andreas Veneris

    Large language models (LLMs) have demonstrated remarkable progress in code generation, but many existing benchmarks are approaching saturation and offer little guarantee on the trustworthiness of the generated programs. To improve visibility into model reasoning on formal correctness, we introduce VerifyThisBench, a new benchmark that evaluates end-to-end pr

  56. Nitin Jha, Abhishek Parakh, Mahadevan Subramaniam

    Quantum-augmented networks aim to use quantum phenomena to improve detection and protection against malicious actors in a classical communication network. This may include multiplexing quantum signals into classical fiber optical channels and incorporating purely quantum links alongside classical links in the network. In such hybrid networks, quantum protoco

  57. Anshu, David Jekel, Therese Basa Landry

    We seek an analog for the quantum permutation group $S_n^+$ of the normalized Hamming distance for permutations. We define three distances on the tracial state space of $C(S_n^+)$ that generalize the $L^1$-Wasserstein distance of probability measures on $S_n$ equipped with the normalized Hamming metric, for which we demonstrate basic metric properties, subad

  58. Javier Cacheiro, Álvaro C Sánchez, Russell Rundle, George B Long

    High-Performance Computing (HPC) systems are the most powerful tools that we currently have to solve complex scientific simulations. Quantum computing (QC) has the potential to enhance HPC systems by accelerating the execution of specific kernels that can be offloaded to a Quantum Processing Unit (QPU), granting them new capabilities, improving the speed of

  59. Yaxuan Yang, Shiyu Wang, Xiaoming Zhai

    Assessing teachers' pedagogical content knowledge (PCK) through performance-based tasks is both time and effort-consuming. While large language models (LLMs) offer new opportunities for efficient automatic scoring, little is known about whether LLMs introduce construct-irrelevant variance (CIV) in ways similar to or different from traditional machine learnin

  60. Yi-Sian Ciou

    The artificial fluid model known as "Schr\"odinger flow" (SF) can represent rotational flow with dissipative effects, and has attracted attention despite its gap from real-world fluid behavior. To address the structural discrepancy arising from the incomplete transition from quantum hydrodynamics to classical fluid dynamics, we propose a variant of the hydro

  61. Guangan Chen, Anh Minh Truong, Hanhe Lin, Michiel Vlaminck

    Novel view synthesis in 360$^\circ$ scenes from extremely sparse input views is essential for applications like virtual reality and augmented reality. This paper presents a novel framework for novel view synthesis in extremely sparse-view cases. As typical structure-from-motion methods are unable to estimate camera poses in extremely sparse-view cases, we ap

  62. Hui Ma, Kai Yang, Yang Jiao

    Network traffic prediction plays a crucial role in intelligent network operation. Traditional prediction methods often rely on centralized training, necessitating the transfer of vast amounts of traffic data to a central server. This approach can lead to latency and privacy concerns. To address these issues, federated learning integrated with differential pr

  63. Lorenzo Bernazzani, Balázs Gulácsi, Guido Burkard

    We investigate in parallel two common pictures used to describe quantum systems interacting with their surrounding environment, i.e., the stochastic Hamiltonian description, where the environment is implicitly included in the fluctuating internal parameters of the system, and the explicit inclusion of the environment via the time-convolutionless projection o

  64. Yu Zhang, Jialei Zhou, Xinchen Li, Qi Zhang

    Current text-to-image diffusion generation typically employs complete-text conditioning. Due to the intricate syntax, diffusion transformers (DiTs) inherently suffer from a comprehension defect of complete-text captions. One-fly complete-text input either overlooks critical semantic details or causes semantic confusion by simultaneously modeling diverse sema

  65. Shiyu Xiang, Tong Zhang, Ronghao Chen

    LLM Agents are becoming central to intelligent systems. However, their deployment raises serious safety concerns. Existing defenses largely rely on "Safety Checks", which struggle to capture the complex semantic risks posed by harmful user inputs or unsafe agent behaviors - creating a significant semantic gap between safety checks and real-world risks. To br

  66. Hossein Zaremehrjerdi, Shreyan Ganguly, Ashlyn Rairdin, Elizabeth Tranel

    Agricultural decision-making involves complex, context-specific reasoning, where choices about crops, practices, and interventions depend heavily on geographic, climatic, and economic conditions. Traditional large language models (LLMs) often fall short in navigating this nuanced problem due to limited reasoning capacity. We hypothesize that recent advances

  67. Felipe Curcio, Pedro Castro, Augusto Fonseca, Rafaela Castro

    With the increasing availability of meteorological data from various sensors, numerical models and reanalysis products, the need for efficient data integration methods has become paramount for improving weather forecasts and hydrometeorological studies. In this work, we propose a data fusion approach for precipitation nowcasting by integrating data from mete

  68. Vivek Gopalakrishnan, Neel Dey, Polina Golland

    Determining the 3D pose of a patient from a limited set of 2D X-ray images is a critical task in interventional settings. While preoperative volumetric imaging (e.g., CT and MRI) provides precise 3D localization and visualization of anatomical targets, these modalities cannot be acquired during procedures, where fast 2D imaging (X-ray) is used instead. To in

  69. Mingyuan Wu, Jingcheng Yang, Jize Jiang, Meitang Li

    Reinforcement Learning Finetuning (RFT) has significantly advanced the reasoning capabilities of large language models (LLMs) by enabling long chains of thought, self-correction, and effective tool use. While recent works attempt to extend RFT to vision-language models (VLMs), these efforts largely produce text-only reasoning conditioned on static image inpu

  70. Rafał Poświata, Marcin Michał Mirończuk, Sławomir Dadas, Małgorzata Grębowiec

    Consumers often face inconsistent product quality, particularly when identical products vary between markets, a situation known as the dual quality problem. To identify and address this issue, automated techniques are needed. This paper explores how natural language processing (NLP) can aid in detecting such discrepancies and presents the full process of dev

  71. João Coelho, Jingjie Ning, Jingyuan He, Kangrui Mao

    Deep research systems represent an emerging class of agentic information retrieval methods that generate comprehensive and well-supported reports to complex queries. However, most existing frameworks rely on dynamic commercial search APIs, which pose reproducibility and transparency challenges in addition to their cost. To address these limitations, we intro

  72. Davin Choo, Billy Jin, Yongho Shin

    Online bipartite matching is a fundamental problem in online optimization, extensively studied both in its integral and fractional forms due to its theoretical significance and practical applications, such as online advertising and resource allocation. Motivated by recent progress in learning-augmented algorithms, we study online bipartite fractional matchin

  73. Stavros Garoufalidis, Matthew Harper, Rinat Kashaev, Ben-Michael Kohli

    Building further on work of Marin and Wagner, we give a cubic braid-type skein theory of the Links--Gould polynomial invariant of oriented links and prove that it can be used to evaluate any oriented link, adding this polynomial to the list of polynomial invariants that can be computed by skein theory. As a consequence, we prove that this skein theory is als

  74. Yi Wang, Junxiao Liu, Shimao Zhang, Jiajun Chen

    Current large-language models (LLMs) typically adopt a fixed reasoning strategy, either simple or complex, for all questions, regardless of their difficulty. This neglect of variation in task and reasoning process complexity leads to an imbalance between performance and efficiency. Existing methods attempt to implement training-free fast-slow thinking system

  75. Mir Sazzat Hossain, Khan Muhammad Bin Asad, Payaswini Saikia, Adrita Khan

    We introduce a novel machine learning dataset tailored for the classification of bent radio active galactic nuclei (AGN) in astronomical observations. Bent radio AGN, distinguished by their curved jet structures, provide critical insights into galaxy cluster dynamics, interactions within the intracluster medium, and the broader physics of AGN. Despite their

  76. Aleksandra E. Nazarova, John M. Cannon, Igor D. Karachentsev, Dmitry I. Makarov

    We describe the results of observations with the 100m Robert C. Byrd Green Bank Telescope (GBT) in the HI line of 105 nearby dwarf galaxies, 60 of which were discovered recently in the DESI Legacy Imaging Surveys. Of 105 objects observed, we detected 77 galaxies with the following median parameters: an HI-flux of 0.69 Jy km/s, a heliocentric velocity of 732

  77. Tao Wang, Ruipeng Zhang, Sicun Gao

    Modern policy gradient algorithms, such as TRPO and PPO, outperform vanilla policy gradient in many RL tasks. Questioning the common belief that enforcing approximate trust regions leads to steady policy improvement in practice, we show that the more critical factor is the enhanced value estimation accuracy from more value update steps in each iteration. To

  78. Jean-Baptiste Camps, Julien Randon-Furling, Ulysse Godreau

    Our knowledge of past cultures relies considerably on written material. For centuries, texts have been copied, altered, then transmitted or lost - eventually, from surviving documents, philologists attempt to reconstruct text phylogenies ("stemmata"), and past written cultures. Nonetheless, fundamental questions on the extent of losses, representativeness of

  79. Kevin Xu, Issei Sato

    Chain-of-Thought (CoT) and Looped Transformers have been shown to empirically improve performance on reasoning tasks and to theoretically enhance expressivity by recursively increasing the number of computational steps. However, their comparative capabilities are still not well understood. In this paper, we provide a formal analysis of their respective stren

  80. Dominik Stempień, Janusz Gajda

    We compare traditional approach of computing logarithmic returns with the fractional differencing method and its tempered extension as methods of data preparation before their usage in advanced machine learning models. Differencing parameters are estimated using multiple techniques. The empirical investigation is conducted on data from four major stock indic

  81. Alaa Dalaq, Muzammil Behzad

    Image segmentation is a fundamental task in computer vision, aimed at partitioning an image into semantically meaningful regions. Referring image segmentation extends this task by using natural language expressions to localize specific objects, requiring effective integration of visual and linguistic information. In this work, we propose SegVLM, a vision-lan

  82. Aida Kostikova, Zhipin Wang, Deidamea Bajri, Ole Pütz

    Large language model (LLM) research has grown rapidly, along with increasing concern about their limitations. In this survey, we conduct a data-driven, semi-automated review of research on limitations of LLMs (LLLMs) from 2022 to early 2025 using a bottom-up approach. From a corpus of 250,000 ACL and arXiv papers, we identify 14,648 relevant papers using key

  83. Chen Shi, Shaoshuai Shi, Kehua Sheng, Bo Zhang

    Data-driven learning has advanced autonomous driving, yet task-specific models struggle with out-of-distribution scenarios due to their narrow optimization objectives and reliance on costly annotated data. We present DriveX, a self-supervised world model that learns generalizable scene dynamics and holistic representations (geometric, semantic, and motion) f

  84. Sourav Ganguly, Kishan Panaganti, Arnob Ghosh, Adam Wierman

    Constrained decision-making is essential for designing safe policies in real-world control systems, yet simulated environments often fail to capture real-world adversities. We consider the problem of learning a policy that will maximize the cumulative reward while satisfying a constraint, even when there is a mismatch between the real model and an accessible

  85. Iñaki Dellibarda Varela, Pablo Romero-Sorozabal, Diego Torricelli, Gabriel Delgado-Oleas

    Self-recognition -- the ability to maintain an internal representation of one's own body within the environment -- underpins intelligent, autonomous behavior. As a foundational component of the minimal self, self-recognition provides the initial substrate from which higher forms of self-awareness may eventually emerge. Recent advances in large language model

  86. Qian Cao, Xiting Wang, Yuzhuo Yuan, Yahui Liu

    Creativity evaluation remains a challenging frontier for large language models (LLMs). Current evaluations heavily rely on inefficient and costly human judgments, hindering progress in enhancing machine creativity. While automated methods exist, ranging from psychological testing to heuristic- or prompting-based approaches, they often lack generalizability o

  87. Debdeep Bhattacharya, Davood Damircheli, Robert P. Lipton

    We present a high-fidelity three dimensional computational framework for simulating the bulk mechanical behavior of granular aggregates composed of deformable brittle grains. Departing from classical discrete element methods (DEM), our approach captures both inter-particle and intra-particle deformation using a nonlocal continuum formulation based on peridyn

  88. Qinsi Wang, Hancheng Ye, Ming-Yu Chung, Yudong Liu

    Vision-Language Models (VLMs) excel across diverse tasks but suffer from high inference costs in time and memory. Token sparsity mitigates inefficiencies in token usage, while neuron sparsity reduces high-dimensional computations, both offering promising solutions to enhance efficiency. Recently, these two sparsity paradigms have evolved largely in parallel,

  89. Jialong Zhou, Lichao Wang, Xiao Yang

    The emergence of large language models (LLMs) enables the development of intelligent agents capable of engaging in complex and multi-turn dialogues. However, multi-agent collaboration faces critical safety challenges, such as hallucination amplification and error injection and propagation. This paper presents GUARDIAN, a unified method for detecting and miti

  90. Aniruddha Mukherjee, Spriha Dubey, Somdyuti Paul

    The rapid advancement of generative AI has enabled the creation of highly photorealistic visual content, offering practical substitutes for real images and videos in scenarios where acquiring real data is difficult or expensive. However, reliably substituting real visual content with AI-generated counterparts requires robust assessment of the perceived realn

  91. Marianna Chantzi, Vassilis Kostopoulos, Spyridon Psarras

    The study analyzes open hole carbon fiber reinforced polymer CFRP laminates modified with electrospun interleaves containing Diels Alder-based self-healing agents. It develops a high-fidelity simulation framework to investigate the quasistatic tensile behavior of these composites. The study uses Hashin's failure criteria to capture intralaminar damage and su

  92. Constantinos Rouvalis, Vassilis Kostopoulos, Spyridon Psarras

    The predictive capabilities of the finite element approach were assessed by comparing simulation outputs with experimental results, including load-displacement trends, damage initiation points, and delamination evolution. This comparison validated the effectiveness of the self-healing interleaves and highlighted the strengths and limitations of the adopted n

  93. Astrid M. Veronig, Karin Dissauer, Bernhard Kliem, Cooper Downs

    Coronal dimmings associated with coronal mass ejections (CME) from the Sun have gained much attention since the late 1990s when they were first observed in high-cadence imagery of the SOHO/EIT and Yohkoh/SXT instruments. They appear as localized sudden decreases of the coronal emission at extreme ultraviolet (EUV) and soft X-ray (SXR) wavelengths, that evolv

  94. Frederik Kunstner, Francis Bach

    Recent works have highlighted optimization difficulties faced by gradient descent in training the first and last layers of transformer-based language models, which are overcome by optimizers such as Adam. These works suggest that the difficulty is linked to the heavy-tailed distribution of words in text data, where the frequency of the $k$th most frequent wo

  95. Efstathios Stratakos, Vassilis Kostopoulos, Spyridon Psarras

    This research aims to develop a method to reduce the time cost and complexity of balloon expandable stent simulations in cardiovascular stenting procedures for Peripheral Artery Disease. The study uses stereoscopic images to construct a cardiovascular stent and validate its mechanical response through Finite Element testing. A 3D model of the curved common f

  96. Chenglong Ma, Yuanfeng Ji, Jin Ye, Zilong Li

    Autoregressive modeling has driven major advances in multimodal AI, yet its application to medical imaging remains constrained by the absence of a unified image tokenizer that simultaneously preserves fine-grained anatomical structures and rich clinical semantics across heterogeneous modalities. Existing approaches jointly optimize image reconstruction and t

  97. Jane Elisa Guimarães, Rafael Nadas, Rayan Alves, Wenjin Zhang

    Contaminations in the formation of two-dimensional heterostructures can hinder or generate desired properties. Recent advancements have highlighted the potential of tip-enhanced Raman spectroscopy (TERS) for studying materials in the 2D semiconductor class. In this work, we investigate the influence of 50-200nm sized nanoprotuberances within a monolayer of M

  98. Fengqi Zhu, Rongzhen Wang, Shen Nie, Xiaolu Zhang

    While Masked Diffusion Models (MDMs), such as LLaDA, present a promising paradigm for language modeling, there has been relatively little effort in aligning these models with human preferences via reinforcement learning. The challenge primarily arises from the high variance in Evidence Lower Bound (ELBO)-based likelihood estimates required for preference opt

  99. Alireza Tabatabaei Mashayekh, Jeremy Witzens

    Light sheet fluorescence microscopy (LSFM) has transformed the way we visualize biological tissues in three dimensions, offering high-resolution imaging while minimizing photo-induced damage to the samples. Recent breakthroughs in tissue-clearing methods have further improved LSFM's capabilities, making it possible to study larger, intact samples in unpreced

  100. Chengbo He, Bochao Zou, Junliang Xing, Jiansheng Chen

    In human-AI collaboration, a central challenge is deciding whether the AI should handle a task, be deferred to a human expert, or be addressed through collaborative effort. Existing Learning to Defer approaches typically make binary choices between AI and humans, neglecting their complementary strengths. They also lack interpretability, a critical property i