May 2025 arXiv papers — page 83
Showing 8,201–8,300 of 24,552 papers
Jeffrey V. Backus, Carolina Figueiredo
Over the past year, the "scalar-scaffolding" formalism has revealed a number of new features of gluon amplitudes. In this paper, we leverage these developments to study two distinct but related questions, linked by the scaffolding statement of gauge invariance. We start by revisiting the soft expansion of gluon amplitudes. The scaffolding picture allows for
Julian Heeck, Mikheil Sokhashvili, Anil Thapa
The conservation of lepton flavor is a prediction of the Standard Model and is still an excellent approximate symmetry despite our observation of neutrino oscillations. Lepton flavor violation by one or two units have been discussed for decades, with several dedicated experiments exploring the vast model landscape but no discoveries so far. Here, we explore
O. Caliskan, M. Uzundag, M. Kilic, F. C. Geronimo
We present extensive follow-up time-series photometry of WD J0049$-$2525, the most massive pulsating white dwarf currently known with $T_{\rm eff} = 13\, 020\,{\rm K}$ and $\log{\it g} = 9.34$ cm s$^{-2}$. The discovery observations detected only two significant pulsation modes. Here, we report the detection of 13 significant pulsation modes ranging from 170
João D. Álvares, Alex Vaño-Viñuales
We present simulations of the Einstein-Maxwell-Klein-Gordon system on compactified hyperboloidal slices. To the best of our knowledge, this is the first time that this setup is evolved with a common formulation like BSSN/Z4. Hyperboloidal slices smoothly reach future null infinity, the only location in spacetime where radiation (such as gravitational waves)
Greg Kaplanek, Alexander Maloney, Jason Pollack, Dylan VanAllen
We give a simple argument that, for a large class of jump operators, the Lindblad evolution can be written as a gradient flow in the space of density operators acting on a Hilbert space of dimension $D$. We give explicit expressions for the (matrix-valued) eigenvectors and eigenvalues of the Lindblad evolution using this formalism. We argue that in many case
A Morphological Model to Separate Resolved-Unresolved Sources in the DESI Legacy Surveys: Application in the LS4 Alert Stream
astro-ph.IMChang Liu, Adam A. Miller, Joshua S. Bloom, Robert A. Knop
Separating resolved and unresolved sources in large imaging surveys is a fundamental step to enable downstream science, such as searching for extragalactic transients in wide-field time-domain surveys. Here we present our method to effectively separate point sources from the resolved, extended sources in the Dark Energy Spectroscopic Instrument (DESI) Legacy
Probing the origins. II. Unravelling lithium depletion and stellar motion: Intrinsic stellar properties drive depletion, not kinematics
astro-ph.SRM. L. L. Dantas, R. Smiljanic, D. Romano, G. Guiglion
In Paper I, we classified a stellar sample from the thin disc with a broad range in metallicity as being churned outward or inward, or blurred/undisturbed. In this paper (Paper II), we delve deeper by analysing our entire metallicity-stratified sample along with their dynamic properties, focusing on the connection between radial migration and Li depletion. W
GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
cs.CVChengqi Duan, Rongyao Fang, Yuqing Wang, Kun Wang
Visual generation models have made remarkable progress in creating realistic images from text prompts, yet struggle with complex prompts that specify multiple objects with precise spatial relationships and attributes. Effective handling of such prompts requires explicit reasoning about the semantic content and spatial layout. We present GoT-R1, a framework t
Sara Ghaboura, Ketan More, Wafa Alghallabi, Omkar Thawakar
As Large Multimodal Models (LMMs) become more capable, there is growing interest in evaluating their reasoning processes alongside their final outputs. However, most benchmarks remain focused on English, overlooking languages with rich linguistic and cultural contexts, such as Arabic. To address this gap, we introduce the Comprehensive Arabic Multimodal Reas
Shilin Yan, Jiaming Han, Joey Tsai, Hongwei Xue
The advent of Large Multimodal Models (LMMs) has significantly enhanced Large Language Models (LLMs) to process and interpret diverse data modalities (e.g., image and video). However, as input complexity increases, particularly with long video sequences, the number of required tokens has grown significantly, leading to quadratically computational costs. This
Let Androids Dream of Electric Sheep: A Human-Inspired Image Implication Understanding and Reasoning Framework
cs.CVChenhao Zhang, Yazhe Niu
Metaphorical comprehension in images remains a critical challenge for AI systems, as existing models struggle to grasp the nuanced cultural, emotional, and contextual implications embedded in visual content. While multimodal large language models (MLLMs) excel in general Visual Question Answer (VQA) tasks, they struggle with a fundamental limitation on image
Kaixuan Fan, Kaituo Feng, Haoming Lyu, Dongzhan Zhou
Recent advances have shown success in eliciting strong reasoning abilities in multimodal large language models (MLLMs) through rule-based reinforcement learning (RL) with outcome rewards. However, this paradigm typically lacks supervision over the thinking process leading to the final outcome. As a result, the model may learn sub-optimal reasoning strategies
Chengzhuo Tong, Ziyu Guo, Renrui Zhang, Wenyu Shan
Recent advancements underscore the significant role of Reinforcement Learning (RL) in enhancing the Chain-of-Thought (CoT) reasoning capabilities of large language models (LLMs). Two prominent RL algorithms, Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO), are central to these developments, showcasing different pros and con
Shuhan Tan, Kairan Dou, Yue Zhao, Philipp Krähenbühl
We introduce RIPT-VLA, a simple and scalable reinforcement-learning-based interactive post-training paradigm that fine-tunes pretrained Vision-Language-Action (VLA) models using only sparse binary success rewards. Existing VLA training pipelines rely heavily on offline expert demonstration data and supervised imitation, limiting their ability to adapt to new
Kevin Lu, Nicky Kriplani, Rohit Gandikota, Minh Pham
In concept erasure, a model is modified to selectively prevent it from generating a target concept. Despite the rapid development of new methods, it remains unclear how thoroughly these approaches remove the target concept from the model. We begin by proposing two conceptual models for the erasure mechanism in diffusion models: (i) interfering with the model
Haoning Wu, Xiao Huang, Yaohui Chen, Ya Zhang
Existing evaluations of multimodal large language models (MLLMs) on spatial intelligence are typically fragmented and limited in scope. In this work, we aim to conduct a holistic assessment of the spatial understanding capabilities of modern MLLMs and propose complementary data-driven and agent-based solutions. Specifically, we make the following contributio
Yan Li, Changyao Tian, Renqiu Xia, Ning Liao
We propose AdapTok, an adaptive temporal causal video tokenizer that can flexibly allocate tokens for different frames based on video content. AdapTok is equipped with a block-wise masking strategy that randomly drops tail tokens of each block during training, and a block causal scorer to predict the reconstruction quality of video frames using different num
Tim Genewein, Li Kevin Wenliang, Jordi Grau-Moya, Anian Ruoss
Prompting is one of the main ways to adapt a pretrained model to target tasks. Besides manually constructing prompts, many prompt optimization methods have been proposed in the literature. Method development is mainly empirically driven, with less emphasis on a conceptual understanding of prompting. In this paper we discuss how optimal prompting can be under
Topological Phase Transitions and Mixed State Order in a Hubbard Quantum Simulator
cond-mat.quant-gasLin Su, Rahul Sahay, Michal Szurek, Alexander Douglas
Topological phase transitions challenge conventional paradigms in many-body physics by separating phases that are locally indistinguishable yet globally distinct. Using a quantum simulator of interacting erbium atoms in an optical lattice, we observe such a transition between one-dimensional crystalline symmetry-protected topological phases (CSPTs). We detec
Jean Pablo Vieira de Mello, Matheus Augusto Alves Cuglieri, Leandro P. de Figueiredo, Fernando Bordignon
Interpreting the mineralogical aspects of rock thin sections is an important task for oil and gas reservoirs evaluation. However, human analysis tend to be subjective and laborious. Technologies like QEMSCAN(R) are designed to automate the mineralogical mapping process, but also suffer from limitations like high monetary costs and time-consuming analysis. Th
Dan Kondo, Takahiro Morimoto, Genta Osaki, Thanaporn Sichanugrist
We propose a novel method to detect axion dark matter based on a topological phenomenon known as the shift current. We make use of the second-order nonlinearity of the shift current by applying a strong oscillating electric field. This field enhances the axion-induced shift current signal and downconverts its frequency to a more accessible range. The nondiss
R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning
cs.CLHuatong Song, Jinhao Jiang, Wenqing Tian, Zhipeng Chen
Large Language Models (LLMs) are powerful but prone to hallucinations due to static knowledge. Retrieval-Augmented Generation (RAG) helps by injecting external information, but current methods often are costly, generalize poorly, or ignore the internal knowledge of the model. In this paper, we introduce R1-Searcher++, a novel framework designed to train LLMs
Jiachen Yao, Abbas Mammadov, Julius Berner, Gavin Kerrigan
We propose a general framework for conditional sampling in PDE-based inverse problems, targeting the recovery of whole solutions from extremely sparse or noisy measurements. This is accomplished by a function-space diffusion model and plug-and-play guidance for conditioning. Our method first trains an unconditional, discretization-agnostic denoising model us
Nanda H. Krishna, Colin Bredenberg, Daniel Levenstein, Blake A. Richards
During periods of quiescence, such as sleep, neural activity in many brain circuits resembles that observed during periods of task engagement. However, the precise conditions under which task-optimized networks can autonomously reactivate the same network states responsible for online behavior is poorly understood. In this study, we develop a mathematical fr
Abdul Hannan, Muhammad Arslan Manzoor, Shah Nawaz, Muhammad Irzam Liaqat
We study the task of learning association between faces and voices, which is gaining interest in the multimodal community lately. These methods suffer from the deliberate crafting of negative mining procedures as well as the reliance on the distant margin parameter. These issues are addressed by learning a joint embedding space in which orthogonality constra
Ming Qian, Bin Tan, Qiuyu Wang, Xianwei Zheng
This paper studies the task of SatStreet-view synthesis, which aims to render photorealistic street-view panorama images and videos given any satellite image and specified camera positions or trajectories. We formulate to learn neural radiance field from paired images captured from satellite and street viewpoints, which comes to be a challenging learning pro
Simmaco Di Lillo
This work investigates the expected number of critical points of random neural networks with different activation functions as the depth increases in the infinite-width limit. Under suitable regularity conditions, we derive precise asymptotic formulas for the expected number of critical points of fixed index and those exceeding a given threshold. Our analysi
Yuan Fang, Shouvik Sur, Yonglong Xie, Qimiao Si
Quantum geometry may enable the development of quantum phases ranging from superconductivity to correlated topological states. One powerful probe of quantum geometry is the nonlinear Hall response which detects Berry curvature dipole in systems with time-reversal invariance and broken inversion symmetry. With broken time-reversal symmetry, this response is a
Jin Jiang, Jianing Wang, Yuchen Yan, Yang Liu
Large Language Models (LLMs) have been shown to achieve breakthrough performance on complex logical reasoning tasks. Nevertheless, most existing research focuses on employing formal language to guide LLMs to derive reliable reasoning paths, while systematic evaluations of these capabilities are still limited. In this paper, we aim to conduct a comprehensive
Rui Ye, Xiangrui Liu, Qimin Wu, Xianghe Pang
LLM-based multi-agent systems (MAS) extend the capabilities of single LLMs by enabling cooperation among multiple specialized agents. However, most existing MAS frameworks rely on a single LLM to drive all agents, constraining the system's intelligence to the limit of that model. This paper explores the paradigm of heterogeneous LLM-driven MAS (X-MAS), where
A Unified Framework for Simultaneous Parameter and Function Discovery in Differential Equations
cs.LGShalev Manor, Mohammad Kohandel
Inverse problems involving differential equations often require identifying unknown parameters or functions from data. Existing approaches, such as Physics-Informed Neural Networks (PINNs), Universal Differential Equations (UDEs) and Universal Physics-Informed Neural Networks (UPINNs), are effective at isolating either parameters or functions but can face ch
DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimization
cs.CLChao Zhang, Xin Shi, Xueqiao Zhang, Yifan Zhu
Recent advances in Emotional Support Conversation (ESC) have improved emotional support generation by fine-tuning Large Language Models (LLMs) via Supervised Fine-Tuning (SFT). However, common psychological errors still persist. While Direct Preference Optimization (DPO) shows promise in reducing such errors through pairwise preference learning, its effectiv
Runyang You, Yongqi Li, Xinyu Lin, Xin Zhang
Large recommender models have extended LLMs as powerful recommenders via encoding or item generation, and recent breakthroughs in LLM reasoning synchronously motivate the exploration of reasoning in recommendation. In this work, we propose R$^2$ec, a unified large recommender model with intrinsic reasoning capability. R$^2$ec introduces a dual-head architect
Ritabrata Biswas, Satyajit Pal
The thermodynamics of black holes (BHs) within the Einstein Maxwell Scalar (EMS) framework, incorporating Barrow entropy and its logarithmic corrections to analyze quantum gravity effects is investigated here. A static, spherically symmetric BH solution is obtained by coupling the scalar field nonminimally to the electromagnetic field through a scalar depend
Guillem Brasó, Aljoša Ošep, Laura Leal-Taixé
Uniform downsampling remains the de facto standard for reducing spatial resolution in vision backbones. In this work, we propose an alternative design built around a content-aware spatial grouping layer, that dynamically assigns tokens to a reduced set based on image boundaries and their semantic content. Stacking our grouping layer across consecutive backbo
PICT -- A Differentiable, GPU-Accelerated Multi-Block PISO Solver for Simulation-Coupled Learning Tasks in Fluid Dynamics
cs.LGAleksandra Franz, Hao Wei, Luca Guastoni, Nils Thuerey
Despite decades of advancements, the simulation of fluids remains one of the most challenging areas of in scientific computing. Supported by the necessity of gradient information in deep learning, differentiable simulators have emerged as an effective tool for optimization and learning in physics simulations. In this work, we present our fluid simulator PICT
Abdul Hannan, Alessio Brutti, Shah Nawaz, Mubashir Noman
Recent advancement in deep learning encouraged developing large automatic speech recognition (ASR) models that achieve promising results while ignoring computational and memory constraints. However, deploying such models on low resource devices is impractical despite of their favorable performance. Existing approaches (pruning, distillation, layer skip etc.)
Runpeng Yu, Xinyin Ma, Xinchao Wang
In this work, we propose Dimple, the first Discrete Diffusion Multimodal Large Language Model (DMLLM). We observe that training with a purely discrete diffusion approach leads to significant training instability, suboptimal performance, and severe length bias issues. To address these challenges, we design a novel training paradigm that combines an initial au
Julia Liebert, Christian Schilling, David A. Mazziotti
We develop a systematic framework for the spin adaptation of the cumulants of p-particle reduced density matrices (RDMs), with explicit constructions for p = 1 to 3. These spin-adapted cumulants enable rigorous treatment of both S_z and S^2 symmetries in quantum systems, providing a foundation for spin-resolved electronic structure methods. We show that comp
Valery V. Ryzhikov
We show slow convergence of weighted ergodic averages for flows and actions of countable amenable groups.
Amartya Chakraborty, Paresh Dashore, Nadia Bathaee, Anmol Jain
Large Language Models (LLMs) have demonstrated impressive capabilities as intelligent agents capable of solving complex problems. However, effective planning in scenarios involving dependencies between API or tool calls-particularly in multi-turn conversations-remains a significant challenge. To address this, we introduce T1, a tool-augmented, multi-domain,
Extremely Simple Multimodal Outlier Synthesis for Out-of-Distribution Detection and Segmentation
cs.CVMoru Liu, Hao Dong, Jessica Kelly, Olga Fink
Out-of-distribution (OOD) detection and segmentation are crucial for deploying machine learning models in safety-critical applications such as autonomous driving and robot-assisted surgery. While prior research has primarily focused on unimodal image data, real-world applications are inherently multimodal, requiring the integration of multiple modalities for
Mingyang Liu, Gabriele Farina, Asuman Ozdaglar
Post-training has demonstrated its importance in enhancing the reasoning capabilities of large language models (LLMs). The primary post-training methods can be categorized into supervised fine-tuning (SFT) and reinforcement fine-tuning (RFT). SFT is efficient and well-suited for small language models, but it may lead to overfitting and limit the reasoning ab
LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position Encoding
cs.CLJunlong Tong, Jinlan Fu, Zixuan Lin, Yingqi Fan
Large Language Models (LLMs) are primarily designed for batch processing. Existing methods for adapting LLMs to streaming rely either on expensive re-encoding or specialized architectures with limited scalability. This work identifies three key mismatches in adapting batch-oriented LLMs to streaming: (1) input-attention, (2) output-attention, and (3) positio
Adib Bazgir, Amir Habibdoust Lafmajani, Yuwen Zhang
Large Language Models (LLMs) show promise in biomedicine but lack true causal understanding, relying instead on correlations. This paper envisions causal LLM agents that integrate multimodal data (text, images, genomics, etc.) and perform intervention-based reasoning to infer cause-and-effect. Addressing this requires overcoming key challenges: designing saf
F. F. Faria
We find that the total entropy of the massive conformal gravity universe is an increasing function of time, and therefore the cosmological model of the theory passes the generalized second law of thermodynamics test.
Phoenix Alpine, Samriddhi Bhatia, Ana M. Botti, Brenda A. Cervantes-Vergara
The Dark matter Nanosatellite Equipped with Skipper Sensors (DarkNESS) deploys a recently developed skipper-CCD architecture with sub-electron readout noise in low Earth orbit (LEO) to investigate potential signatures of dark matter (DM). The mission addresses two interaction channels: electron recoils from strongly interacting sub-GeV DM and X-rays produced
Dong Li, Wenqi Zhong, Wei Yu, Yingwei Pan
Video virtual try-on aims to seamlessly dress a subject in a video with a specific garment. The primary challenge involves preserving the visual authenticity of the garment while dynamically adapting to the pose and physique of the subject. While existing methods have predominantly focused on image-based virtual try-on, extending these techniques directly to
Zhenkun Li, Lingyao Li, Shuhang Lin, Yongfeng Zhang
Single-agent LLMs hit hard limits--finite context, role overload, and brittle domain transfer. Conventional multi-agent fixes soften those edges yet expose fresh pains: ill-posed decompositions, fuzzy contracts, and verification overhead that blunts the gains. We therefore present Know-The-Ropes (KtR), a framework that converts domain priors into an algorith
Weizhi Tang, Yixuan Li, Chris Sypherd, Elizabeth Polgreen
Grammar plays a critical role in natural language processing and text/code generation by enabling the definition of syntax, the creation of parsers, and guiding structured outputs. Although large language models (LLMs) demonstrate impressive capabilities across domains, their ability to infer and generate grammars has not yet been thoroughly explored. In thi
Siqi Wan, Jingwen Chen, Yingwei Pan, Ting Yao
Diffusion models have shown preliminary success in virtual try-on (VTON) task. The typical dual-branch architecture comprises two UNets for implicit garment deformation and synthesized image generation respectively, and has emerged as the recipe for VTON task. Nevertheless, the problem remains challenging to preserve the shape and every detail of the given g
Yurui Qian, Qi Cai, Yingwei Pan, Ting Yao
Contemporary diffusion models show remarkable capability in text-to-image generation, while still being limited to restricted resolutions (e.g., 1,024 X 1,024). Recent advances enable tuning-free higher-resolution image generation by recycling pre-trained diffusion models and extending them via regional denoising or dilated sampling/convolutions. However, th
Yaxin Du, Yuzhu Cai, Yifan Zhou, Cheng Wang
Large Language Models (LLMs) have shown strong capability in diverse software engineering tasks. However, feature-driven development, a highly prevalent real-world task that involves developing new functionalities for large, existing codebases, remains underexplored. We therefore introduce SWE-Dev, the first large-scale dataset (with 14,000 training and 500
Zongyan Han, Jiale Cao, Shuo Chen, Tong Wang
Open-Vocabulary Segmentation (OVS) has drawn increasing attention for its capacity to generalize segmentation beyond predefined categories. However, existing methods typically predict segmentation masks with simple forward inference, lacking explicit reasoning and interpretability. This makes it challenging for OVS model to distinguish similar categories in
Rishanth Rajendhran, Amir Zadeh, Matthew Sarte, Chuan Li
Metrics like FactScore and VeriScore that evaluate long-form factuality operate by decomposing an input response into atomic claims and then individually verifying each claim. While effective and interpretable, these methods incur numerous LLM calls and can take upwards of 100 seconds to evaluate a single response, limiting their practicality in large-scale
Tianduo Wang, Lu Xu, Wei Lu, Shanbo Cheng
Recent advances in Automatic Speech Recognition (ASR) have been largely fueled by massive speech corpora. However, extending coverage to diverse languages with limited resources remains a formidable challenge. This paper introduces Speech Back-Translation, a scalable pipeline that improves multilingual ASR models by converting large-scale text corpora into s
Himangi Mittal, Peiye Zhuang, Hsin-Ying Lee, Shubham Tulsiani
We propose UniPhy, a common latent-conditioned neural constitutive model that can encode the physical properties of diverse materials. At inference UniPhy allows `inverse simulation' i.e. inferring material properties by optimizing the scene-specific latent to match the available observations via differentiable simulation. In contrast to existing methods tha
Christopher Criscitiello, Jungbin Kim
Geodesic convexity (g-convexity) is a natural generalization of convexity to Riemannian manifolds. However, g-convexity lacks many desirable properties satisfied by Euclidean convexity. For instance, the natural notions of half-spaces and affine functions are themselves not g-convex. Moreover, recent studies have shown that the oracle complexity of geodesica
Boce Hu, Dian Wang, David Klee, Heng Tian
Equivariant models have recently been shown to improve the data efficiency of diffusion policy by a significant margin. However, prior work that explored this direction focused primarily on point cloud inputs generated by multiple cameras fixed in the workspace. This type of point cloud input is not compatible with the now-common setting where the primary in
Ahmed Heakl, Gustavo Bertolo Stahl, Sarim Hashmi, Seung Hun Eddie Han
Cross-architecture GPU code transpilation is essential for unlocking low-level hardware portability, yet no scalable solution exists. We introduce CASS, the first dataset and model suite for source- and assembly-level GPU translation (CUDA <--> HIP, SASS <--> RDNA3). CASS contains 60k verified host-device code pairs, enabling learning-based translation acros
Simulating Time Dependent and Nonlinear Classical Oscillators through Nonlinear Schr\"odingerization
quant-phAbhinav Muraleedharan, Nathan Wiebe
We present quantum algorithms for simulating the dynamics of a broad class of classical oscillator systems containing $2^n$ coupled oscillators (Eg: $2^n$ masses coupled by springs), including those with time-dependent forces, time-varying stiffness matrices, and weak nonlinear interactions. This generalization of the Harmonic oscillator simulation algorithm
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
cs.IRNandan Thakur, Crystina Zhang, Xueguang Ma, Jimmy Lin
Training robust retrieval and reranker models typically relies on large-scale retrieval datasets; for example, the BGE collection contains 1.6 million query-passage pairs sourced from various data sources. However, we find that certain datasets can negatively impact model effectiveness -- pruning 8 out of 15 datasets from the BGE collection, reduces the trai
Modeling Inequality in Complex Networks of Strategic Agents using Iterative Game-Theoretic Transactions
cs.GTMayank Kejriwal, Yuesheng Luo
Transactions are an important aspect of human social life, and represent dynamic flow of information, intangible values, such as trust, as well as monetary and social capital. Although much research has been conducted on the nature of transactions in fields ranging from the social sciences to game theory, the systemic effects of different types of agents tra
BP-Seg: A graphical model approach to unsupervised and non-contiguous text segmentation using belief propagation
cs.CLFengyi Li, Kayhan Behdin, Natesh Pillai, Xiaofeng Wang
Text segmentation based on the semantic meaning of sentences is a fundamental task with broad utility in many downstream applications. In this paper, we propose a graphical model-based unsupervised learning approach, named BP-Seg for efficient text segmentation. Our method not only considers local coherence, capturing the intuition that adjacent sentences ar
Suhao Yu, Haojin Wang, Juncheng Wu, Luyang Luo
Real-world clinical practice demands multi-image comparative reasoning, yet current medical benchmarks remain limited to single-frame interpretation. We present MedFrameQA, the first benchmark explicitly designed to test multi-image medical VQA through educationally-validated diagnostic sequences. To construct this dataset, we develop a scalable pipeline tha
Jonas Bayer, Marco David
We present a universal construction of Diophantine equations with bounded complexity in Isabelle/HOL. This is a formalization of our own work in number theory. Hilbert's Tenth Problem was answered negatively by Yuri Matiyasevich, who showed that there is no general algorithm to decide whether an arbitrary Diophantine equation has a solution. However, the pro
Benjamin S. Savino, Amirreza Rouhi, Wen Wu
Turbulent boundary layers over riblets subjected to adverse pressure gradients (APGs) are investigated by direct numerical simulation. Multiple APG strengths and riblet sizes are examined, permitting evaluation of drag modification by riblets, and associated physical mechanisms, in various regimes established for zero-pressure-gradient (ZPG) riblet flows. Th
Julien Froustey
Collisional flavor instabilities, driven by differing neutrino and antineutrino reaction rates, are expected to occur in dense astrophysical environments like supernovae and neutron star mergers, but have yet to be incorporated in large-scale simulations. We derive analytical expressions for the asymptotic state resulting from a homogeneous and isotropic ins
David Zywina
For any quadratic extension $L/K$ of number fields, we prove that there are infinitely many elliptic curves $E$ over $K$ so that the abelian groups $E(K)$ and $E(L)$ both have rank $1$. In particular, there are infinitely many elliptic curves of rank $1$ over any number field. This result generalizes theorems of Koymans-Pagano and Alp\"oge-Bhargava-Ho-Shnidm
Impact, Causation and Prediction of Socio-Academic and Economic Factors in Exam-centric Student Evaluation Measures using Machine Learning and Causal Analysis
cs.LGMd. Biplob Hosen, Sabbir Ahmed, Bushra Akter, Mehrin Anannya
Understanding socio-academic and economic factors influencing students' performance is crucial for effective educational interventions. This study employs several machine learning techniques and causal analysis to predict and elucidate the impacts of these factors on academic performance. We constructed a hypothetical causal graph and collected data from 1,0
Alessandro Favero, Antonio Sclocchi, Matthieu Wyart
Diffusion probabilistic models have become a cornerstone of modern generative AI, yet the mechanisms underlying their generalization remain poorly understood. In fact, if these models were perfectly minimizing their training loss, they would just generate data belonging to their training set, i.e., memorize, as empirically found in the overparameterized regi
André Pedroso Kowacs
We apply the characterization of global hypoellipticity for $G$-invariant operators on homogeneous vector bundles obtained by Cardona and Kowacs [J. Pseudo-Differ. Oper. Appl. 16, 23 (2025)] to obtain a necessary and sufficient condition for an arbitrary system of left-invariant operators on a compact Lie group to be globally hypoelliptic, providing a full p
Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models
cs.CRJunjie Xiong, Changjia Zhu, Shuhang Lin, Chong Zhang
Large Language Models (LLMs) are increasingly equipped with capabilities of real-time web search and integrated with protocols like Model Context Protocol (MCP). This extension could introduce new security vulnerabilities. We present a systematic investigation of LLM vulnerabilities to hidden adversarial prompts through malicious font injection in external r
Daniil Gurgurov, Michal Gregor, Josef van Genabith, Simon Ostermann
In this paper, we combine two-step knowledge distillation, structured pruning, truncation, and vocabulary trimming for extremely compressing multilingual encoder-only language models for low-resource languages. Our novel approach systematically combines existing techniques and takes them to the extreme, reducing layer depth, feed-forward hidden size, and int
Roger Casals, Kenton Ke
We study the boundedness of a mutation class for quivers with real weights. The main result is a characterization of bounded mutation classes for real quivers of rank 3.
Cracking Aegis: An Adversarial LLM-based Game for Raising Awareness of Vulnerabilities in Privacy Protection
cs.HCJiaying Fu, Yiyang Lu, Zehua Yang, Fiona Nah
Traditional methods for raising awareness of privacy protection often fail to engage users or provide hands-on insights into how privacy vulnerabilities are exploited. To address this, we incorporate an adversarial mechanic in the design of the dialogue-based serious game Cracking Aegis. Leveraging LLMs to simulate natural interactions, the game challenges p
Young Sang Choi, Vincent Jeanselme, Pierre Elias, Shalmali Joshi
Multimodal learning is of continued interest in artificial intelligence-based applications, motivated by the potential information gain from combining different data modalities. However, modalities observed in the source environment may differ from the modalities observed in the target environment due to multiple factors, including cost, hardware failure, or
FrontierCO: Real-World and Large-Scale Evaluation of Machine Learning Solvers for Combinatorial Optimization
cs.LGShengyu Feng, Weiwei Sun, Shanda Li, Ameet Talwalkar
Machine learning (ML) has shown promise for tackling combinatorial optimization (CO), but much of the reported progress relies on small-scale, synthetic benchmarks that fail to capture real-world structure and scale. A core limitation is that ML methods are typically trained and evaluated on synthetic instance generators, leaving open how they perform on irr
Santiago Berrezueta-Guzman, Andrei Koshelev, Stefan Wagner
Photogrammetry is transforming digital content creation by enabling the rapid conversion of real-world objects into highly detailed 3D models. This paper evaluates the role of RealityCapture, a GPU-accelerated photogrammetry tool, in game development of Virtual Reality (VR). We assess its efficiency, reconstruction accuracy, and integration with Unreal Engin
Adnan Oomerjee, Zafeirios Fountas, Haitham Bou-Ammar, Jun Wang
Transformer LLMs have been shown to exhibit strong reasoning ability that scales with inference-time compute, most prominently through token-space "thinking" chains of thought. A growing line of work pushes extra computation into the model's latent space, which we term Auxiliary Latent-Space Computation (ALSC). Existing ALSC methods largely fall into three b
Gautam Bharali, Rumpa Masanta
In this paper, we explore some connections between Kobayashi geometry and the Dirichlet problem for the complex Monge--Amp\`ere equation. Among the results we obtain through these connections are: $(i)$~a theorem on the continuous extension up to $\partial{D}$ of a proper holomorphic map $F: D\longrightarrow \Omega$ between domains with $\dim_{\mathbb{C}}(D)
Dhruv Devulapalli, Chao Yin, Andrew Y. Guo, Eddie Schoute
To implement arbitrary quantum circuits in architectures with restricted interactions, one may effectively simulate all-to-all connectivity by routing quantum information. We consider the entanglement dynamics and routing between two regions only connected through an intermediate "bottleneck" region with few qubits. In such systems, where the entanglement ra
Csaba Dékány, Stefan Balauca, Robin Staab, Dimitar I. Dimitrov
Despite recent efforts in Large Language Model (LLM) safety and alignment, current adversarial attacks on frontier LLMs can still consistently force harmful generations. Although adversarial training has been widely studied and shown to significantly improve the robustness of traditional machine learning models, its strengths and weaknesses in the context of
Sanjana Chalavadi, Andrei Pastor, Terry Leitch
This study analyzes tract-level real estate ownership patterns in New York State (NYS) and New York City (NYC) to uncover racial disparities. We use an advanced race/ethnicity imputation model (LSTM+Geo with XGBoost filtering, validated at 89.2% accuracy) to compare the predicted racial composition of property owners to the resident population from census da
Adam Chudecki
A special class of (complex) para-Hermite Einstein spaces is analyzed. It is well-known that the self-dual Weyl tensor in para-Hermite Einstein spaces is of the Petrov-Penrose type [D]. In what follows we assume that the anti-self-dual Weyl tensor is algebraically degenerate. It is equivalent to the existence of an anti-self-dual congruence of null strings w
Yunjia Qi, Hao Peng, Xiaozhi Wang, Amy Xin
Large Language Models (LLMs) have demonstrated advanced capabilities in real-world agentic applications. Growing research efforts aim to develop LLM-based agents to address practical demands, introducing a new challenge: agentic scenarios often involve lengthy instructions with complex constraints, such as extended system prompts and detailed tool specificat
R. Xu, B. N. J. Persson
We present a study of sliding friction for rigid triangular steel sliders on soft rubber substrates under both lubricated and dry conditions. For rubber surfaces lubricated with a thin film of silicone oil, the measured sliding friction at room temperature agrees well with theoretical predictions obtained from a viscoelastic model originally developed for ro
FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records
cs.LGVincent Jeanselme, Zilin Jing, Aparajita Kashyap, Chao Pang
Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (iii) poor transportability. Despite methodological advances in structured electronic health record (EHR) foundation models, no systematic benchmark has validated whether these model
Center-symmetric Landau gauge, the deconfinement transition and the gluon propagator as seen in lattice QCD
hep-latDuifje Maria van Egmond, Orlando Oliveira, Urko Reinosa, Julien Serreau
We address the lattice computation of the gluon propagator in the center-symmetric Landau gauge. After discussing a proper lattice implementation of the center-symmetric Landau gauge, we compare the lattice data with analytical results, and we identify various signatures of center symmetry breaking.
Adrian Saldanha, Adam Peichl, Wim Michiels, Tomáš Vyhlídal
We present a methodology for designing a dynamic controller with delayed output feedback for achieving non-collocated vibration suppression with a focus on the multi-frequency case. To synthesize the delay-based controller, we first remodel the system of equations as a delay-differential algebraic equation (DDAE) in such a way that existing tools for design
InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification
cs.AIInternAgent Team, Bo Zhang, Shiyang Feng, Xiangchao Yan
Artificial Intelligence (AI) is accelerating the transformation of scientific research paradigms, not only enhancing research efficiency but also driving innovation. We introduce InternAgent, a unified closed-loop multi-agent framework to conduct Autonomous Scientific Research (ASR) across various scientific research fields, enabling researchers to tackle co
Noah Amsel, Tyler Chen, Feyza Duman Keles, Diana Halikias
We present a randomized algorithm for producing a quasi-optimal hierarchically semi-separable (HSS) approximation to an $N\times N$ matrix $A$ using only matrix-vector products with $A$ and $A^T$. We prove that, using $O(k \log(N/k))$ matrix-vector products and ${O}(N k^2 \log(N/k))$ additional runtime, the algorithm returns an HSS matrix $B$ with rank-$k$ b
Yizhuo Chen, Tianchen Wang, You Lyu, Yanlan Hu
We present SPAR, a framework for self-supervised placement-aware representation learning in distributed sensing. Distributed sensing spans applications where multiple spatially distributed and multimodal sensors jointly observe an environment, from vehicle monitoring to human activity recognition and earthquake localization. A central challenge shared by thi
Mostafaali Ayubirad, Madiha Akbar, Hamid R. Ossareh
This paper addresses the challenge of pressure constraint violations in water electrolysis systems operating under dynamic power conditions, a problem common to both Proton Exchange Membrane and alkaline technologies. To investigate this issue, a control-oriented model of an alkaline electrolyzer is developed, capturing key pressure and flow dynamics. To man
Yepeng Liu, Xuandong Zhao, Christopher Kruegel, Dawn Song
The growing use of large language models (LLMs) for sensitive applications has highlighted the need for effective watermarking techniques to ensure the provenance and accountability of AI-generated text. However, most existing watermarking methods require access to the decoding process, limiting their applicability in real-world settings. One illustrative ex
Zebin You, Shen Nie, Xiaolu Zhang, Jun Hu
In this work, we introduce LLaDA-V, a purely diffusion-based Multimodal Large Language Model (MLLM) that integrates visual instruction tuning with masked diffusion models, representing a departure from the autoregressive paradigms dominant in current multimodal approaches. Built upon LLaDA, a representative large language diffusion model, LLaDA-V incorporate
Noah Amsel, David Persson, Christopher Musco, Robert M. Gower
Computing the polar decomposition and the related matrix sign function has been a well-studied problem in numerical analysis for decades. Recently, it has emerged as an important subroutine within the Muon optimizer for training deep neural networks. However, the requirements of this application differ sharply from classical settings: deep learning demands G
Matthew Zent, Digory Smith, Simon Woodhead
Personally identifiable information (PII) anonymization is a high-stakes task that poses a barrier to many open-science data sharing initiatives. While PII identification has made large strides in recent years, in practice, error thresholds and the recall/precision trade-off still limit the uptake of these anonymization pipelines. We present PIIvot, a lighte
A. F. Morais, M. C. Araújo, T. T. Saraiva, J. Furtado
We studied a Lorentz-violating inspired Ginzburg-Landau model for superconductivity where we considered a CPT-odd contribution given by $(k_{AF})^{\mu}$, also known as the Carroll-Field-Jackiw term. In the static limit of the equations, we could find a pair of modified Ginzburg-Landau equations. Furthermore, these equations were reduced to the London equatio
Pietro Klausner, Marco Antonelli, Francesca Gulminelli
We perform a Bayesian analysis of the neutron star (NS) equation of state (EoS) based on a wide set of Skyrme functionals, derived from previous nuclear physics inferences. The novelty of this approach lies in starting from the full multidimensional posterior distribution of nuclear matter parameters, consistent with a comprehensive set of static and dynamic