Skip to content

May 2025 arXiv papers — page 73

Showing 7,2017,300 of 24,552 papers

  1. Zizhang Li, Hong-Xing Yu, Wei Liu, Yin Yang

    WonderPlay is a novel framework integrating physics simulation with video generation for generating action-conditioned dynamic 3D scenes from a single image. While prior works are restricted to rigid body or simple elastic dynamics, WonderPlay features a hybrid generative simulator to synthesize a wide range of 3D dynamics. The hybrid generative simulator fi

  2. Mengru Wang, Ziwen Xu, Shengyu Mao, Shumin Deng

    Precise control over language model generation is vital for ensuring both safety and reliability. Although prompt engineering and steering are commonly used to intervene in model behaviors, the vast number of parameters in models often results in highly intertwined internal representations. This interdependency can limit control precision and sometimes lead

  3. Nic Fishman, Gokul Gowri, Peng Yin, Jonathan Gootenberg

    Many real-world problems require reasoning across multiple scales, demanding models which operate not on single data points, but on entire distributions. We introduce generative distribution embeddings (GDE), a framework that lifts autoencoders to the space of distributions. In GDEs, an encoder acts on sets of samples, and the decoder is replaced by a genera

  4. Mathew J. Koretsky, Maya Willey, Owen Bianchi, Chelsea X. Alvarado

    Biomedical researchers increasingly rely on large-scale structured databases for complex analytical tasks. However, current text-to-SQL systems often struggle to map qualitative scientific questions into executable SQL, particularly when implicit domain reasoning is required. We introduce BiomedSQL, the first benchmark explicitly designed to evaluate scienti

  5. Aradhye Agarwal, Ayan Sengupta, Tanmoy Chakraborty

    Test-time scaling (TTS), which involves dynamic allocation of compute during inference, offers a promising way to improve reasoning in large language models. While existing TTS methods work well, they often rely on long decoding paths or require a large number of samples to be generated, increasing the token usage and inference latency. We observe the surpri

  6. Owen Bianchi, Mathew J. Koretsky, Maya Willey, Chelsea X. Alvarado

    Large language models (LLMs) face significant challenges with needle-in-ahaystack tasks, where relevant information ("the needle") must be drawn from a large pool of irrelevant context ("the haystack"). Previous studies have highlighted positional bias and distractor quantity as critical factors affecting model performance, yet the influence of gold context

  7. Wojciech G. Stark, Connor L. Box, Matthias Sachs, Nils Hertl

    Dissociative chemisorption is a key process in hydrogen-metal surface chemistry, where nonadiabatic effects due to low-lying electron-hole-pair excitations may affect reaction outcomes. Molecular dynamics with electronic friction simulations can capture weak nonadiabatic effects at metal surfaces, but require as input energy landscapes and electronic frictio

  8. Mona Azadkia, Pouya Roudaki

    We introduce a novel measure of dependence that captures the extent to which a random variable $Y$ is determined by a random vector $X$. The measure equals zero precisely when $Y$ and $X$ are independent, and it attains one exactly when $Y$ is almost surely a measurable function of $X$. We further extend this framework to define a measure of conditional depe

  9. Mohamed Swailem, Ulrich Dobramysl, Ruslan Mukhamadiarov, Uwe C. Täuber

    We provide an overview of Monte Carlo algorithms based on Markovian stochastic dynamics of interacting and reacting many-particle systems not in thermal equilibrium. These agent-based simulations are an effective way of introducing students to current research without requiring much prior knowledge or experience. By starting from the direct visualization of

  10. Zoltan Dencs, Vera Dobos, Zsolt Regaly

    Of the few thousand discovered exoplanets, a significant number orbit in the habitable zone of their star. Many of them are gas giants lacking a rocky surface and solid water reservoirs necessary for life as we know it. The search for habitable environments may extend to the moons of these giant planets. No confirmed exomoon discoveries have been made as of

  11. Prithvi Raj Datla, Luheng Zhao, Wen Wei Ho, Natalie Klco

    Lattice gauge theories (LGTs) provide a framework for describing dynamical systems ranging from nuclei to materials. LGTs that host concatenated conservation laws can exhibit Hilbert space fragmentation, where each subspace may be labeled by a conserved quantity with nonlocal operator support. It is expected that nonlocal conservation laws will not impede th

  12. Hongyun Wang, Shannon E. Foley, Hong Zhou

    We investigate the temperature evolution in the three-dimensional skin tissue exposed to a millimeter-wave electromagnetic beam that is not necessarily perpendicular to the skin surface. This study examines the effect of the beam's incident angle. The incident angle influences the thermal heating in two aspects: (i) the beam spot projected onto the skin is e

  13. Junfeng Wu, Dongliang Luo, Weizhi Zhao, Zhihao Xie

    In this work, we reveal the limitations of visual tokenizers and VAEs in preserving fine-grained features, and propose a benchmark to evaluate reconstruction performance for two challenging visual contents: text and face. Visual tokenizers and VAEs have significantly advanced visual generation and multimodal modeling by providing more efficient compressed or

  14. Burcu Kilic, Alper Ahmetoglu, Emre Ugur

    Discovering symbolic representations for skills is essential for abstract reasoning and efficient planning in robotics. Previous neuro-symbolic robotic studies mostly focused on discovering perceptual symbolic categories given a pre-defined action repertoire and generating plans with given action symbols. A truly developmental robotic system, on the other ha

  15. Taskin Mehereen, Sourav Saha, Intesar Jawad Jaigirdar, Chanwook Park

    The ability to accurately model interatomic interactions in large-scale systems is fundamental to understanding a wide range of physical and chemical phenomena, from drug-protein binding to the behavior of next-generation materials. While machine learning interatomic potentials (MLIPs) have made it possible to achieve ab initio-level accuracy at significantl

  16. Aline B. Trench, João Paulo C. Moura, Caio Machado Fernandes, Mauro C. Santos

    The oxygen reduction reaction (ORR) via the 2-electron mechanism is an efficient way to produce hydrogen peroxide (H2O2) under mild conditions. This study examines the modification of Vulcan XC72 carbon with fluorine (F)-doped niobium oxide (Nb2O5) nanoparticles at varying molar ratios (0, 0.005, 0.01, 0.02). The F-doped Nb2O5 nanoparticles were synthesized

  17. Gordon Dai, Yunze Xiao

    This position paper argues that the theoretical inconsistency often observed among Responsible AI (RAI) metrics, such as differing fairness definitions or tradeoffs between accuracy and privacy, should be embraced as a valuable feature rather than a flaw to be eliminated. We contend that navigating these inconsistencies, by treating metrics as divergent obje

  18. Dan Kondo, Hitoshi Murayama, Bea Noether

    Gauge theories can be solved exactly slightly away from the supersymmetric (SUSY) limit softly broken by anomaly mediation when the size of SUSY breaking is much smaller than the dynamical scale ($m \ll \Lambda$). We show empirical evidence that the near-SUSY limit is continuously connected to the non-SUSY limit ($m \gg \Lambda$) in $\mathrm{SU}(N_c)$ gauge

  19. Amit Kumar Kundu, Vaishnavi S Patil, Joseph Jaja

    The open set recognition (OSR) problem aims to identify test samples from novel semantic classes that are not part of the training classes, a task that is crucial in many practical scenarios. However, the existing OSR methods use a constant scaling factor (the temperature) to the logits before applying a loss function, which hinders the model from exploring

  20. Mykola Trokhymovych, Lydia Pintscher, Ricardo Baeza-Yates, Diego Saez-Trumper

    We introduce a next-generation vandalism detection system for Wikidata, one of the largest open-source structured knowledge bases on the Web. Wikidata is highly complex: its items incorporate an ever-expanding universe of factual triples and multilingual texts. While edits can alter both structured and textual content, our approach converts all edits into a

  21. Kazem Faghih, Wenxiao Wang, Yize Cheng, Siddhant Bharti

    Large language models (LLMs) can now access a wide range of external tools, thanks to the Model Context Protocol (MCP). This greatly expands their abilities as various agents. However, LLMs rely entirely on the text descriptions of tools to decide which ones to use--a process that is surprisingly fragile. In this work, we expose a vulnerability in prevalent

  22. Nitin Jha, Abhishek Parakh, Mahadevan Subramaniam

    Secure quantum networks are a bedrock requirement for developing a future quantum internet. However, quantum channels are susceptible to channel noise that introduce errors in the transmitted data. The traditional approach to providing error correction typically encapsulates the message in an error correction code after encryption. Such separate processes in

  23. Dingqiang Ye, Chao Fan, Zhanbo Huang, Chengwen Luo

    Large vision models (LVM) based gait recognition has achieved impressive performance. However, existing LVM-based approaches may overemphasize gait priors while neglecting the intrinsic value of LVM itself, particularly the rich, distinct representations across its multi-layers. To adequately unlock LVM's potential, this work investigates the impact of layer

  24. Jonas A. Actor, Graham Harper, Ben Southworth, Eric C. Cyr

    Multilayer perceptrons (MLPs) are a workhorse machine learning architecture, used in a variety of modern deep learning frameworks. However, recently Kolmogorov-Arnold Networks (KANs) have become increasingly popular due to their success on a range of problems, particularly for scientific machine learning tasks. In this paper, we exploit the relationship betw

  25. Charles D. Coleman

    Measuring the accuracy of cross-sectional predictions is a subjective problem. Generally, this problem is avoided. In contrast, this paper confronts subjectivity up front by eliciting an impartial decision-maker's preferences. These preferences are embedded into an axiomatically-derived loss function, one of the simplest version of which is described. The pa

  26. Yan Ma, Linge Du, Xuyang Shen, Shaoxiang Chen

    Reinforcement learning (RL) is becoming an important direction for post-training vision-language models (VLMs), but public training methodologies for unified multimodal RL remain much less mature, especially for heterogeneous reasoning and perception-heavy tasks. We propose V-Triune, a Visual Triple Unified Reinforcement Learning methodology for unified mult

  27. Chau Minh Pham, Jenna Russell, Dzung Pham, Mohit Iyyer

    We introduce Frankentexts, a long-form narrative generation paradigm that treats an LLM as a composer of existing texts rather than as an author. Given a writing prompt and thousands of randomly sampled human-written snippets, the model is asked to produce a narrative under the extreme constraint that most tokens (e.g., 90%) must be copied verbatim from the

  28. Rustam Arabov, Nikita Rybin, Victor Demin, Mikhail Polovinkin

    We have investigated the effect of interlayer twist angle on lattice thermal conductivity (LTC) and band gap renormalization in boron nitride and carbon Moir\'e diamanes. Moment tensor potentials were used for calculating energies and forces of interatomic interactions. The methods based on the solution of Boltzmann transport equation (BTE) for phonons and t

  29. Lorenz Wolf, Robert Kirk, Mirco Musolesi

    Reinforcement learning from human feedback (RLHF) is a widely used method for aligning large language models with human preferences. However, RLHF often suffers from reward model overoptimisation, in which models overfit to the reward function, resulting in non-generalisable policies that exploit the idiosyncrasies and peculiarities of the reward function. A

  30. Alan Arazi, Eilam Shapira, Roi Reichart

    While deep learning has achieved remarkable success across many domains, it has historically underperformed on tabular learning tasks, which remain dominated by gradient boosting decision trees. However, recent advancements are paving the way for Tabular Foundation Models, which can leverage real-world knowledge and generalize across diverse datasets, partic

  31. Liuke Lyu, Deeksha Chandorkar, Samarth Kapoor, So Takei

    Quantum spin liquids (QSLs) give rise to exotic emergent particles by weaving intricate entanglement patterns in the underlying electrons. Bipartite measures between subregions can detect the presence of anyons, but little is known about the full entanglement structure of QSLs. Here, we study the multiparty entanglement of QSLs via entanglement microscopy. W

  32. Poojah Ganesan, Rajat Aayush Jha, Dan Roth, Vivek Gupta

    Recent advances in large language models (LLMs) have greatly improved Text-to-SQL performance for single-table queries. But, it remains challenging in multi-table databases due to complex schema and relational operations. Existing methods often struggle with retrieving the right tables and columns, generating accurate JOINs and UNIONs, and generalizing acros

  33. Danyang Zhang, Situo Zhang, Ziyue Yang, Zichen Zhu

    LLM-based (Large Language Model) GUI (Graphical User Interface) agents can potentially reshape our daily lives significantly. However, current LLM-based GUI agents suffer from the scarcity of high-quality training data owing to the difficulties of trajectory collection and reward annotation. Existing works have been exploring LLMs to collect trajectories for

  34. Jiongran Wu, Jiahao Liu, Dongsheng Li, Guangping Zhang

    Large language models (LLMs) have demonstrated exceptional performance in understanding and generating semantic patterns, making them promising candidates for sequential recommendation tasks. However, when combined with conventional recommendation models (CRMs), LLMs often face challenges related to high inference costs and static knowledge transfer methods.

  35. Aidan Gleich, Eric Laber, Alexander Volfovsky

    Many interventions, such as vaccines in clinical trials or coupons in online marketplaces, must be assigned sequentially without full knowledge of their effects. Multi-armed bandit algorithms have proven successful in such settings. However, standard independence assumptions fail when the treatment status of one individual impacts the outcomes of others, a p

  36. Frigyes Samuel Racz, John Milton, Juan Luis Cabrera, Gábor Csukly

    Aperiodic neural activity has been the subject of intense research interest lately as it could reflect on the cortical excitation/inhibition ratio, which is suspected to be affected in numerous clinical conditions. This phenomenon is characterized via the aperiodic scaling exponent $\beta$, equal to the spectral slope following log-log transformation of powe

  37. Kunal Sawarkar, Shivam R. Solanki, Abhilasha Mangal

    Retrieval-Augmented Generation (RAG) struggles with domain-specific enterprise datasets, often isolated behind firewalls and rich in complex, specialized terminology unseen by LLMs during pre-training. Semantic variability across domains like medicine, networking, or law hampers RAG's context precision, while fine-tuning solutions are costly, slow, and lack

  38. Huayu Chen, Kaiwen Zheng, Qinsheng Zhang, Ganqu Cui

    Reinforcement Learning (RL) has played a central role in the recent surge of LLMs' math abilities by enabling self-improvement through binary verifier signals. In contrast, Supervised Learning (SL) is rarely considered for such verification-driven training, largely due to its heavy reliance on reference answers and inability to reflect on mistakes. In this w

  39. Jacob Hansen, Wei Lin, Junmo Kang, Muhammad Jehanzeb Mirza

    Visual Instruction Tuning (VisIT) data, commonly available as human-assistant conversations with images interleaved in the human turns, are currently the most widespread vehicle for aligning strong LLMs to understand visual inputs, converting them to strong LMMs. While many VisIT datasets are available, most are constructed using ad-hoc techniques developed

  40. Richard Cole, Pranav Jangir

    In the facility location problem, the task is to place one or more facilities so as to minimize the sum of the agent costs for accessing their nearest facility. Heretofore, in the strategic version, agent locations have been assumed to be private, while their cost measures have been public and identical. For the most part, the cost measure has been the dista

  41. Yuqi Wang, Sirui Wang, Shiman Zhang, Kexue Fu

    Performance artforms like Peking opera face transmission challenges due to the extensive passive listening required to understand their nuance. To create engaging forms of experiencing auditory Intangible Cultural Heritage (ICH), we designed a spatial interaction-based segmented-audio (SISA) Virtual Reality system that transforms passive ICH experiences into

  42. Cheng-Yen Yang, Hsiang-Wei Huang, Pyong-Kun Kim, Chien-Kai Kuo

    We present an effective approach for adapting the Segment Anything Model 2 (SAM2) to the Visual Object Tracking (VOT) task. Our method leverages the powerful pre-trained capabilities of SAM2 and incorporates several key techniques to enhance its performance in VOT applications. By combining SAM2 with our proposed optimizations, we achieved a first place AUC

  43. Zinuo Li, Xian Zhang, Yongxin Guo, Mohammed Bennamoun

    Humans naturally understand moments in a video by integrating visual and auditory cues. For example, localizing a scene in the video like "A scientist passionately speaks on wildlife conservation as dramatic orchestral music plays, with the audience nodding and applauding" requires simultaneous processing of visual, audio, and speech signals. However, existi

  44. Juliane Haug, Fabian Wunder

    We revisit the recently published analytic results for unpolarized and polarized semi-inclusive deep inelastic scattering (SIDIS) at next-to-next-to-leading order (NNLO) in QCD. These expressions for the hard scattering coefficients contain case distinctions in the kinematic $(x,z)$ plane splitting the analytic result in four regions. By re-expressing the co

  45. Cristina Ana-Maria Anghel

    We construct geometrically two universal link invariants: universal ADO invariant and universal Jones invariant, as limits of invariants given by graded intersections in configuration spaces. More specifically, for a fixed level $\mathscr N$, we define new link invariants: ``$\mathscr N^{th}$ Unified Jones invariant'' and ``$\mathscr N^{th}$ Unified Alexande

  46. Yichi Zhang, Zhihao Duan, Yuning Huang, Fengqing Zhu

    As learned image compression (LIC) methods become increasingly computationally demanding, enhancing their training efficiency is crucial. This paper takes a step forward in accelerating the training of LIC methods by modeling the neural training dynamics. We first propose a Sensitivity-aware True and Dummy Embedding Training mechanism (STDET) that clusters L

  47. Varun Ajith, Anindya Pal, Saumik Bhattacharya, Sayantari Ghosh

    Nanomaterial research is becoming a vital area for energy, medicine, and materials science, and accurate analysis of the nanoparticle topology is essential to determine their properties. Unfortunately, the lack of high-quality annotated datasets drastically hinders the creation of strong segmentation models for nanoscale imaging. To alleviate this problem, w

  48. Yusuf Yildiz, Goran Nenadic, Meghna Jani, David A. Jenkins

    Objective: Large language models (LLMs) are attracting increasing interest in healthcare. This commentary evaluates the potential of LLMs to improve clinical prediction models (CPMs) for diagnostic and prognostic tasks, with a focus on their ability to process longitudinal electronic health record (EHR) data. Findings: LLMs show promise in handling multimoda

  49. Lisheng Huang, Yichen Liu, Jinhao Jiang, Rongxiang Zhang

    Recent advances in web-augmented large language models (LLMs) have exhibited strong performance in complex reasoning tasks, yet these capabilities are mostly locked in proprietary systems with opaque architectures. In this work, we propose \textbf{ManuSearch}, a transparent and modular multi-agent framework designed to democratize deep search for LLMs. ManuS

  50. Roy Elkayam

    This study presents a novel approach for decomposing urban water demand patterns using Skewed Gaussian Distributions (SGD) to derive behavioral insights and support operational planning. Hourly demand profiles contain critical information for both long-term infrastructure design and daily operations, influencing network pressures, water quality, energy consu

  51. Asher Auel, Jack Petok

    We define the zeta function of a noncommutative K3 surface over a finite field, an invariant under Fourier-Mukai equivalence that can be used to define point counts in this noncommutative setting. These point counts can be negative, and can be used as an obstruction to geometricity. In particular, we study the K3 category associated to a cubic fourfold over

  52. Alessandro R. Mazza, Jia Shi, Gabriel A. Vázquez-Lizardi, Sangsoo Kim

    The epitaxial synthesis of high-quality 2D layered materials is an essential driver of both fundamental physics studies and technological applications. Bi$_2$Se$_3$, a prototypical 2D layered topological insulator, is sensitive to defects imparted during the growth, either thermodynamically or due to the film-substrate interaction. In this study, it is shown

  53. Congren Dai, Huichi Zhou, Jiahao Huang, Zhenxuan Zhang

    Online Continual Learning (OCL) involves sequentially arriving data and is particularly challenged by catastrophic forgetting, which significantly impairs model performance. To address this issue, we introduce a novel framework, Online Dynamic Expandable Dual Memory (ODEDM), that integrates a short-term memory for fast memory and a long-term memory structure

  54. Yukin Zhang, Qi Dong, Kemu Xu

    Why do language models from different architecture families respond so differently to the same perturbation? We argue that the answer is not scale, but \emph{how architecture shapes information compression}. Analyzing eight Transformer models (7B--70B parameters) from the Llama and Qwen families, we show that every model spontaneously develops discrete funct

  55. Omar Madany, Benjamin Kincaid, Aqsa Shaikh, Elizabeth Morningstar

    We present a new set of correlation-consistent effective core potentials (ccECPs) for selected heavy $s$, $p$, $d$, and $f$-block elements significant in materials science and chemistry (Rb, Sr, Cs, Ba, In, Sb, Pb, Ru, Cd, La, Ce, and Eu). The ccECPs are designed using minimal Gaussian parameterization to achieve smooth and bounded potentials. They are expre

  56. Yuxin Liu, M. Amin Rahimian, Kiran Garimella

    WhatsApp, a platform with more than two billion global users, plays a crucial role in digital communication, but also serves as a vector for harmful content such as misinformation, hate speech, and political propaganda. This study examines the dynamics of harmful message dissemination in WhatsApp groups, with a focus on their structural characteristics. Usin

  57. Joey Hong, Anca Dragan, Sergey Levine

    Large language models (LLMs) excel in tasks like question answering and dialogue, but complex tasks requiring interaction, such as negotiation and persuasion, require additional long-horizon reasoning and planning. Reinforcement learning (RL) fine-tuning can enable such planning in principle, but suffers from drawbacks that hinder scalability. In particular,

  58. Chun Tong Lei, Zhongliang Guo, Hon Chung Lee, Minh Quoc Duong

    Adversarial attacks have become a well-explored domain, frequently serving as evaluation baselines for model robustness. Among these, black-box attacks based on transferability have received significant attention due to their practical applicability in real-world scenarios. Traditional black-box methods have generally focused on improving the optimization fr

  59. Ziqiao Peng, Yanbo Fan, Haoyu Wu, Xuan Wang

    In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive conversation, which leads to unnatural interactions and awkward transitions. To address this issue, we propose a new task --

  60. Jackson K. Wilt, Natalie M. Larson, Jennifer A. Lewis

    The rapid design and fabrication of soft robotic matter is of growing interest for shape morphing, actuation, and wearable devices. Here, we report a facile fabrication method for creating soft robotic materials with embedded pneumatics that exhibit programmable shape morphing behavior. Using rotational multi-material 3D printing, asymmetrical core-shell fil

  61. John W. Patty, Elizabeth Maggie Penn

    We demonstrate that the set of cost distributions under which the optimal strategy for maximizing compliance (or more generally, effort) in a binary choice environment is identical to the optimal strategy for maximizing the accuracy of the reward (minimizing Type-I and Type-II errors) is finitely shy (Anderson and Zame (2001) in the space of all smooth param

  62. Jitendra Kethepalli, Andrew Urilyon, Tridib Sadhu, Jacopo De Nardis

    Ballistic Macroscopic Fluctuation Theory (BMFT) captures the evolution of fluctuations and correlations in systems where transport is strictly ballistic. We show that, for \emph{generic integrable models}, BMFT can be constructed through a direct mapping onto ensembles of classical or quantum point particles. This mapping generalises the well-known correspon

  63. Weizhou Shen, Chenliang Li, Fanqi Wan, Shengyi Liao

    This technical report presents QwenLong-CPRS, a context compression framework designed for explicit long-context optimization, addressing prohibitive computation overhead during the prefill stage and the "lost in the middle" performance degradation of large language models (LLMs) during long sequence processing. Implemented through a novel dynamic context op

  64. Xinran Gu, Kaifeng Lyu, Jiazheng Li, Jingzhao Zhang

    Large Language Models (LLMs) are typically trained on data mixtures: most data come from web scrapes, while a small portion is curated from high-quality sources with dense domain-specific knowledge. In this paper, we show that when training LLMs on such data mixtures, knowledge acquisition from knowledge-dense datasets, unlike training exclusively on knowled

  65. Vytautas Abramavicius, Evan Philip, Kaonan Micadei, Charles Moussa

    Parameter shift rules are instrumental for derivatives estimation in a wide range of quantum algorithms, especially in the context of Quantum Machine Learning. Application of single-gap parameter shift rule is often not possible in algorithms running on noisy intermediate-scale quantum (NISQ) hardware due to noise effects and interaction between device qubit

  66. Allen O. Scheie, Sabrina J. Li, Stephen D. Wilson, Daniel A. Rehn

    A growing list of Ce-based magnets have shown an extra and heretofore unexplained crystal electric field (CEF) mode at high energies. We describe a process whereby an optical phonon can produce a split CEF mode well above the phonon energy. We use density functional theory and point-charge model calculations to estimate the phonon distortions and coupling to

  67. Andrea Giuseppe Di Francesco, Maria Sofia Bucarelli, Franco Maria Nardini, Raffaele Perego

    Early-exit mechanisms allow deep neural networks to stop inference once prediction confidence is high, reducing latency and energy on easy inputs while retaining full-depth accuracy on harder ones. Similarly, adding early exit mechanisms to Graph Neural Networks (GNNs), the go-to models for graph-structured data, allows for dynamic trading depth for confiden

  68. Hyungyung Lee, Geon Choi, Jung-Oh Lee, Hangyul Yoon

    Recent progress in Large Vision-Language Models (LVLMs) has enabled promising applications in medical tasks, such as report generation and visual question answering. However, existing benchmarks focus mainly on the final diagnostic answer, offering limited insight into whether models engage in clinically meaningful reasoning. To address this, we present CheX

  69. Muzhi Dai, Shixuan Liu, Qingyi Si

    The success of Deepseek-R1 has drawn the LLM community's attention to reinforcement learning (RL) methods like GRPO. However, such rule-based 0/1 outcome reward methods lack the capability to regulate the intermediate reasoning processes during chain-of-thought (CoT) generation, leading to severe overthinking phenomena. In response, recent studies have desig

  70. Santiago Jaraba, Sachiko Kuroyanagi, Qiuyue Liang, Meng-Xiang Lin

    Astrometry, the precise measurement of stellar positions and velocities, offers a promising approach to probing the low-frequency stochastic gravitational wave background (SGWB). Notably, astrometric vector sky maps are sensitive to parity-violating SGWB signals, which cannot be distinguished using pulsar timing array observations in an isotropic SGWB. We pr

  71. E. D. Emtsova, A. N. Petrov, A. V. Toporensky

    This paper brings a methodological character where we present a comprehensive formalism for constructing conserved quantities in the Teleparallel Equivalent of General Relativity (TEGR) and Symmetric Teleparallel Equivalent of General Relativity (STEGR). It was developed in series of our earlier works and, here, we unite it into a complete form. By employing

  72. Quentin Clark, Florian Shkurti

    In policy learning, stitching and compositional generalization refer to the extent to which the policy is able to piece together sub-trajectories of data it is trained on to generate new and diverse behaviours. While stitching has been identified as a significant strength of offline reinforcement learning, recent generative behavioural cloning (BC) methods h

  73. Georgios Kementzidis, Erin Wong, John Nicholson, Ruichen Xu

    The techniques of data-driven backmapping from coarse-grained (CG) to fine-grained (FG) representation often struggle with accuracy, unstable training, and physical realism, especially when applied to complex systems such as proteins. In this work, we introduce a novel iterative framework by using conditional Variational Autoencoders and graph-based neural n

  74. Adam D. Cobb, Susmit Jha

    Recent work on backpropagation-free learning has shown that it is possible to use forward-mode automatic differentiation (AD) to perform optimization on differentiable models. Forward-mode AD requires sampling a tangent vector for each forward pass of a model. The result is the model evaluation with the directional derivative along the tangent. In this paper

  75. Chunlin Gong, Yin Wang, Jingru Li, Hanleran Zhang

    This paper presents AFD-STA Net, a neural framework integrating adaptive filtering and spatiotemporal dynamics learning for predicting high-dimensional chaotic systems governed by partial differential equations. The architecture combines: 1) An adaptive exponential smoothing module with position-aware decay coefficients for robust attractor reconstruction, 2

  76. Ismail Erbas, Ferhat Demirkiran, Karthik Swaminathan, Naigang Wang

    Fluorescence LiDAR (FLiDAR), a Light Detection and Ranging (LiDAR) technology employed for distance and depth estimation across medical, automotive, and other fields, encounters significant computational challenges in scattering media. The complex nature of the acquired FLiDAR signal, particularly in such environments, makes isolating photon time-of-flight (

  77. Xiaoyi Zhang, Zhaoyang Jia, Zongyu Guo, Jiahao Li

    Long-form video understanding presents significant challenges due to extensive temporal-spatial complexity and the difficulty of question answering under such extended contexts. While Large Language Models (LLMs) have demonstrated considerable advancements in video analysis capabilities and long context handling, they continue to exhibit limitations when pro

  78. Junhao Chen, Mingjin Chen, Jianjin Xu, Xiang Li

    Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange positions under noisy control signals. We address this gap with DanceTogether, the first end-to-end diffusion framework that turns a single reference image plus independent pose-mask streams into long, photorealistic

  79. Wei Wang, Xiaoyu Ou, Zhihan Ren, Waqas Bin Abbas

    In this paper, we focus on the energy efficiency (EE) optimization and analysis of reconfigurable intelligent surface (RIS)-assisted multiuser downlink near-field communications. Specifically, we conduct a comprehensive study on several key factors affecting EE performance, including the number of RIS elements, the types of reconfigurable elements, reconfigu

  80. Yong Wan, Holly A. Holman, Charles Hansen

    Laser scanning microscopy enables the acquisition of 3D data in biomedical research. A fundamental challenge in visualizing 3D data is that common flat-panel displays, being 2D in nature, cannot faithfully reproduce light fields. Recent years have witnessed the development of various 3D display technologies. These technologies generally fall into two categor

  81. Ada Polizzi, Michael Fucilla, Alessandro Papa

    We present a general formula for the amplitude of forward exclusive hadronic processes in the semihard regime of perturbative Quantum Chromodynamics (QCD), by means of the {\em next-to-leading order} eigenfunctions of the Balitsky-Fadin-Kuraev-Lipatov (BFKL) kernel, as constructed by Chirilli and Kovchegov. We discuss some formal subtleties in the check of c

  82. Emile Lavaut

    The Deep Underground Neutrino Experiment (DUNE) is a next-generation long-baseline neutrino experiment. In addition to GeV-scale oscillation measurements ($\delta_{CP}$, $\theta_{23}$ octant, mass ordering), DUNE features a low-energy (MeV-scale) program targeting solar, supernova burst (SNB), and Diffuse Supernova Background (DSNB) neutrinos. Accurate recon

  83. Georg Diez, Nele Dethloff, Gerhard Stock

    Dimensionality reduction represents a crucial step in extracting meaningful insights from Molecular Dynamics (MD) simulations. Conventional approaches, including linear methods such as principal component analysis as well as various autoencoder architectures, typically operate under the assumption of independent and identically distributed data, disregarding

  84. Jia-Nan Li, Jian Guan, Wei Wu, Rui Yan

    Large language models (LLMs) have demonstrated significant success in complex reasoning tasks such as math and coding. In contrast to these tasks where deductive reasoning predominates, inductive reasoning-the ability to derive general rules from incomplete evidence, remains underexplored. This paper investigates extended inductive reasoning in LLMs through

  85. Balbeer Singh

    Jets are extended multipartonic systems and serve as a powerful tool for investigating the dynamics of emergent phenomena driven by many body QCD interactions. In heavy ion collisions, starting from their production during the perturbative hard scattering event in the initial stages of the collision to non-perturbative hadronization they interact with the va

  86. Francesco Di Lauro, Luca Ferretti

    We include complex connectivity structures and heterogeneity in models of multilayer networks or multilayer hypergraphs growing by preferential attachment. We consider the most generic connectivity structure, where the probability of acquiring a new hyperlink depends linearly on the vector of hyperdegrees of the node across all layers, as well as on the laye

  87. Satya Narayana Cheetirala, Ganesh Raut, Dhavalkumar Patel, Fabio Sanatana

    Long text classification is challenging for Large Language Models (LLMs) due to token limits and high computational costs. This study explores whether a Retrieval Augmented Generation (RAG) approach using only the most relevant text segments can match the performance of processing entire clinical notes with large context LLMs. We begin by splitting clinical

  88. Min Hun Lee, Martyn Zhe Yu Tok

    Despite the growing promise of artificial intelligence (AI) in supporting decision-making across domains, fostering appropriate human reliance on AI remains a critical challenge. In this paper, we investigate the utility of exploring distance-based uncertainty scores for task delegation to AI and describe how these scores can be visualized through embedding

  89. Zeen Song, Wenwen Qiang, Siyu Zhao, Changwen Zheng

    External test-time reasoning enhances large language models (LLMs) by decoupling generation and selection. At inference time, the model generates multiple reasoning paths, and an auxiliary process reward model (PRM) is used to score and select the best one. A central challenge in this setting is test-time compute optimality (TCO), i.e., how to maximize answe

  90. Victor Boone

    In this paper, we present a learning algorithm that achieves asymptotically optimal regret for Markov decision processes in average reward under a communicating assumption. That is, given a communicating Markov decision process $M$, our algorithm has regret $K(M) \log(T) + \mathrm{o}(\log(T))$ where $T$ is the number of learning steps and $K(M)$ is the best

  91. Niall Donlon, Romina Gaburro

    We consider the inverse boundary value problem of the simultaneous determination of the coefficients $\sigma$ and $q$ of the equation $-\mbox{div}(\sigma \nabla u)+qu = 0$ from knowledge of the so-called Neumann-to-Dirichlet map, given locally on a non-empty curved portion $\Sigma$ of the boundary $\partial \Omega$ of a domain $\Omega \subset \mathbb{R}^n$,

  92. Renaud Gueroult, Shreekrishna K. Tripathi, Jia Han, Patrick Pribyl

    Because of the speed of light compared to material motion, the dragging of light is difficult to observe under laboratory conditions. Here we report on the first observation of image rotation, i. e. a dragging by the medium of the wave's transverse structure, of Alfv\'en waves in plasmas. Exploiting the naturally slow group velocity of these waves, significa

  93. José Correa, Vasilis Livanos, Dana Pizarro, Victor Verdugo

    Posted price mechanisms are prevalent in allocating goods within online marketplaces due to their simplicity and practical efficiency. We explore a fundamental scenario where buyers' valuations are independent and identically distributed, focusing specifically on the allocation of a single unit. Inspired by the rapid growth and scalability of modern online m

  94. Yumeng Zhang, Shruti Atul Mali, Danial Khan, Sina Amirrajab

    Objectives Accurate MRI-based identification of extramural vascular invasion (EVI) and mesorectal fascia invasion (MFI) is crucial for risk-stratified rectal cancer treatment. However, subjective visual assessment and inter-institutional variability limit diagnostic consistency. This study developed and externally evaluated a multi-centre, foundation model-d

  95. Keenan E. Avers, Jarryd A. Horn, Ram Kumar, Shanta R. Saha

    A very fundamental property of both weakly and strongly interacting materials is the nature of its magnetic response. In this work we detail the growth of crystals of the quasicrystal approximant Fe$_4$Al$_{13}$ with an Al flux solvent method. We characterize our samples using electrical transport and heat capacity, yielding results consistent with a simple

  96. Monirul Islam Mahmud

    ZeroML is a new generation programming language for AutoML to drive the ML pipeline in a compiled and multi-paradigm way, with a pure functional core. Meeting the shortcomings introduced by Python, R, or Julia such as slow-running time, brittle pipelines or high dependency cost ZeroML brings the Microservices-based architecture adding the modular, reusable p

  97. Kunhang Li, Jason Naradowsky, Yansong Feng, Yusuke Miyao

    We explore the human motion knowledge of Large Language Models (LLMs) through 3D avatar control. Given a motion instruction, we prompt LLMs to first generate a high-level movement plan with consecutive steps (High-level Planning), then specify body part positions in each step (Low-level Planning), which we linearly interpolate into avatar animations. Using 2

  98. Wei-Ling Hsu, Yu-Chien Tang, An-Zi Yen

    The increasing reliance on Large Language Models (LLMs) across various domains extends to education, where students progressively use generative AI as a tool for learning. While prior work has examined LLMs' mathematical ability, their reliability in grading authentic student problem-solving processes and delivering effective feedback remains underexplored.

  99. Nao Moriyama

    We discuss the relative log minimal model theory for log surfaces in the analytic setting. More precisely, we show that the minimal model program, the abundance theorem, and the finite generation of log canonical rings hold for log pairs of complex surfaces which are projective over complex analytic varieties.

  100. Jon Merladet Urigüen, Ashot Minasyan

    A group $G$ has property (VRC) if every cyclic subgroup is a virtual retract. This property is stable under many standard group-theoretic constructions and is enjoyed by all virtually special groups (in the sense of Haglund and Wise). In this paper we study property (VRC) for fundamental groups of finite graphs of groups. Our main criterion shows that the fu