May 2025 arXiv papers — page 67
Showing 6,601–6,700 of 24,552 papers
Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning
cs.IRJinzheng Li, Sibo Ju, Yanzhou Su, Hongguang Li
Existing large language models (LLMs) driven search agents typically rely on prompt engineering to decouple the user queries into search plans, limiting their effectiveness in complex scenarios requiring reasoning. Furthermore, they suffer from excessive token consumption due to Python-based search plan representations and inadequate integration of multimedi
Wenlong Deng, Yi Ren, Muchen Li, Danica J. Sutherland
Reinforcement learning (RL) has become popular in enhancing the reasoning capabilities of large language models (LLMs), with Group Relative Policy Optimization (GRPO) emerging as a widely used algorithm in recent systems. Despite GRPO's widespread adoption, we identify a previously unrecognized phenomenon we term Lazy Likelihood Displacement (LLD), wherein t
Kai Mei, Xi Zhu, Hang Gao, Shuhang Lin
We present AIOS 1.0, a novel platform designed to advance computer-use agent (CUA) capabilities through environmental contextualization. While existing approaches primarily focus on building more powerful agent frameworks or enhancing agent models, we identify a fundamental limitation: the semantic disconnect between how language models understand the world
Junyan Liu, Ziyun Chen, Kun Wang, Haipeng Luo
We study the Pandora's Box problem in an online learning setting with semi-bandit feedback. In each round, the learner sequentially pays to open up to $n$ boxes with unknown reward distributions, observes rewards upon opening, and decides when to stop. The utility of the learner is the maximum observed reward minus the cumulative cost of opened boxes, and th
Cesar Ayala, Antonio Pineda
We give the most up-to-date determinations of the normalization of the leading renormalons of the pole mass, the singlet static potential, the octet static potential, and the gluelump energy. They read $Z^{\rm MS}_m=-Z^{\rm MS}_{V_s}/2=\{0.604(17),0.551(20)\}$, $Z^{\rm MS}_{V_o}=\{0.136(8),0.121(13)\}$, and $Z^{\rm MS}_A=\{-1.343(36),-1.224(43)\}$, for $n_f=
Mikhail Ershov, Matthew C. B. Zaremsky
An automorphism of the free group $F_n$ is called pure symmetric if it sends each generator to a conjugate of itself. The group $\mathrm{PSA}_n$ of all pure symmetric automorphisms and its quotient $\mathrm{PSO}_n$ by the group of inner automorphisms are called the McCool groups. In this paper we prove that every BNSR-invariant $\Sigma^m$ of a McCool group i
Nicholas M. Boffi, Michael S. Albergo, Eric Vanden-Eijnden
Flow-based generative models achieve state-of-the-art sample quality, but require the expensive solution of a differential equation at inference time. Flow map models, commonly known as consistency models, encompass many recent efforts to improve inference-time efficiency by learning the solution operator of this differential equation. Yet despite their prom
FlatAttention: Dataflow and Fabric Collectives Co-Optimization for Efficient Multi-Head Attention on Tile-Based Many-PE Accelerators
cs.ARChi Zhang, Luca Colagrande, Renzo Andri, Thomas Benz
Multi-Head Attention (MHA) is a critical computational kernel in transformer-based AI models. Emerging scalable tile-based accelerator architectures integrate increasing numbers of tightly-packed processing elements (PEs) with tensor units. MHA dataflow mapping is crucial for achieving high utilization of the available units. We propose FlatAttention, a new
Mihai Prunescu, Lorenzo Sauras-Altuzarra, Joseph M. Shunia
We show that the class of Kalm\'ar elementary functions can be inductively generated from the addition, the integer remainder, and the base-two exponentiation, hence improving previous results by Marchenkov and Mazzanti. We also prove that the substitution basis defined by these three operations is minimal. Furthermore, we discuss alternative substitution ba
Libin Lan, Yanxin Li, Xiaojuan Liu, Juan Zhou
Accurate medical image segmentation allows for the precise delineation of anatomical structures and pathological regions, which is essential for treatment planning, surgical navigation, and disease monitoring. Both CNN-based and Transformer-based methods have achieved remarkable success in medical image segmentation tasks. However, CNN-based methods struggle
Shijue Huang, Hongru Wang, Wanjun Zhong, Zhaochen Su
Modern large reasoning models demonstrate impressive problem-solving capabilities by employing sophisticated reasoning strategies. However, they often struggle to balance efficiency and effectiveness, frequently generating unnecessarily lengthy reasoning chains for simple problems. In this work, we propose AdaCtrl, a novel framework to support both difficult
Yair Neuman, Yochai Cohen
Traditional interpretations of probability, whether frequentist or subjective, make no reference to the concept of energy. In this paper, we propose that assigning hypothetical energy levels to the outcomes of a random variable can yield improved probability estimates. We apply this Boltzmann-informed approach to the context of sports betting and analyze fiv
Samit Ghosh, Arjun Paul
Let $X$ be an irreducible smooth complex projective variety. Let $G$ be a linear algebraic group over $\mathbb{C}$. We define the notion of Lie algebroid valued connection on holomorphic principal $G$--bundles on $X$, and study their basic properties under extension and reduction of structure group. Finally we investigate criterions for existence of a Lie al
Weijun Kong, Qing Wang
Momentum space entanglement of four fermion field theory is calculated from the Wilsonian effective action pertubatively using replica trick, local terms in low energy effective action are proved to be non-relevant pertubatively and nonlocal terms are the only source of entanglement between different momentum modes. The final result again can be represented
Matthew Sotoudeh
Bespoke data structure operations are common in real-world C code. We identify one common subclass, monotonic data structure traversals (MDSTs), that iterate monotonically through the structure. For example, strlen iterates from start to end of a character array until a null byte is found, and a binary search tree insert iterates from the tree root towards a
High-order Equivariant Flow Matching for Density Functional Theory Hamiltonian Prediction
physics.comp-phSeongsu Kim, Nayoung Kim, Dongwoo Kim, Sungsoo Ahn
Density functional theory (DFT) is a fundamental method for simulating quantum chemical properties, but it remains expensive due to the iterative self-consistent field (SCF) process required to solve the Kohn-Sham equations. Recently, deep learning methods are gaining attention as a way to bypass this step by directly predicting the Hamiltonian. However, the
Yiqing Shen, Chenjia Li, Fei Xiong, Jeong-O Jeong
Reasoning Segmentation (RS) aims to delineate objects based on implicit text queries, the interpretation of which requires reasoning and knowledge integration. Unlike the traditional formulation of segmentation problems that relies on fixed semantic categories or explicit prompting, RS bridges the gap between visual perception and human-like reasoning capabi
Lorenzo Gavassino
A new first-order theory of relativistic dissipation has been recently proposed, where viscous effects are incorporated using the traditional Navier-Stokes framework. Its main novelty is the avoidance of dynamical instabilities by allowing different observers to use equations that are not related by exact Lorentz transformations. In this work, we explore the
Omer Ege, Mustafa Cagal, Kemal Bicakci
As electronic signatures (e-signatures) become increasingly integral to secure digital transactions, understanding their usability and security perception from an end-user perspective has become crucial. This study empirically evaluates and compares two major e-signature systems -- token-based and remote signatures -- through a controlled user experience stu
Entanglement dynamics of Multi-Level Atoms embedded in Photonic Crystals: Leveraging Resonant Dipole-Dipole Interactions and Quantum Interference
quant-phNancy Ghangas, Shubhrangshu Dasgupta
We present a comprehensive investigation of entanglement dynamics in multi-level V-type atomic systems embedded within photonic crystals. We mainly focus on the synergistic roles of resonant dipole-dipole interactions and quantum interference through analytical modeling and numerical simulations using the Schrodinger equation. Key findings reveal that resona
Ye Sun, Hao Zhang, Henghui Ding, Tiehua Zhang
Achieving fine-grained spatio-temporal understanding in videos remains a major challenge for current Video Large Multimodal Models (Video LMMs). Addressing this challenge requires mastering two core capabilities: video referring understanding, which captures the semantics of video regions, and video grounding, which segments object regions based on natural l
Vincent Bouchard, Asia Matthews
Contemporary anarchism centers around three tenets: (1) a constant challenge of and resistance to all forms of domination, (2) so-called "prefigurative politics", in which all decisions are made in a manner that is consistent with a set of non-hierarchical values such as equality, decentralization and voluntary cooperation, (3) a focus on diversity and open-
Philipp L. Kinon, Riccardo Morandin, Philipp Schulze
Discrete gradient methods are a powerful tool for the time discretization of dynamical systems, since they are structure-preserving regardless of the form of the total energy. In this work, we discuss the application of discrete gradient methods to the system class of nonlinear port-Hamiltonian differential-algebraic equations - as they emerge from the port-
Wenhao Sun, Rong-Cheng Tu, Yifu Ding, Zhao Jin
Video diffusion transformers have achieved remarkable progress in high-quality video generation, but remain computationally expensive due to the quadratic complexity of attention over high-dimensional video sequences. Recent acceleration methods enhance the efficiency by exploiting the local sparsity of attention scores; yet they often struggle with accelera
Jiaming Ji, Wenqi Chen, Kaile Wang, Donghai Hong
Modern large language models rely on chain-of-thought (CoT) reasoning to achieve impressive performance, yet the same mechanism can amplify deceptive alignment, situations in which a model appears aligned while covertly pursuing misaligned goals. Existing safety pipelines treat deception as a black-box output to be filtered post-hoc, leaving the model free t
Nam Hoang Thanh, Trung Pham Duy, Lam Bui Thu
Machine learning (ML) has been developed to detect malware in recent years. Most researchers focused their efforts on improving the detection performance but ignored the robustness of the ML models. In addition, many machine learning algorithms are very vulnerable to intentional attacks. To solve these problems, adversarial malware examples are generated by
Zhongtian Zheng, Tao Huang, Haozhe Su, Xueqi Ma
Hair cards remain a widely used representation for hair modeling in real-time applications, offering a practical trade-off between visual fidelity, memory usage, and performance. However, generating high-quality hair card models remains a challenging and labor-intensive task. This work presents an automated pipeline for converting strand-based hair models in
Ammar Younas
Contemporary discussions in AI ethics often treat culture as a source of normative divergence that needs to be accommodated, tolerated, or managed due to its resistance to universal standards. This paper offers an alternative vision through the concept of "Cultural Co-Genesis of AI Ethics." Rather than viewing culture as a boundary or container of isolated m
Valeriy G. Bardakov, Tatyana A. Kozlovskaya, Matvei N. Zonov
We define the Cayley graph and its growth function for multivalued groups. We prove that if we change a finite set of generators of multivalued group, or change the starting point, we get an equivalent growth function. We prove that if we take a virtually nilpotent group and construct a coset group with respect a finite group of authomorphisms, then this mul
A validated coupled three-dimensional hydrodynamic and spectral wind-wave model for the western north Atlantic Ocean
physics.ao-phMaria Venolia, Reza Marsooli, Jaime R. Calzada
Wind-wave and ocean current interactions affect critical coastal and oceanic processes, yet modeling these interactions presents significant challenges. The western North Atlantic Ocean provides an ideal test environment for coupled hydrodynamics and wind wave models, thanks to its energetic surface currents such as the Gulf Stream. This study evaluates a hi
Martina Di Cesare
The fourth observing run (O4) of Advanced LIGO, Virgo, and KAGRA has started in May 2023 and is planned to continue until October 2025. On behalf of the LVK Collaboration, I will cover two topics: Status of the O4 run and latest non-CBC results. Status of the O4 run. The focus will be on detectors' performance and online searches/alerts, drawing on publicly
Nurali Akramov, Karim Rakhimov
In this work, we prove that the complement of the Brjuno set in $\mathbb{C}^n$ has zero $C_\sigma$-capacity with respect to the kernel $k_\sigma(z,\xi)=\|z-\xi\|^{-2n+2}|\log{\|z-\xi\||^{\sigma}}$ for any $\sigma>n$. In particular, it follows that it has zero $h_\delta$-Hausdorff measure with respect to the $h_\delta(t)=t^{2n-2}|\log{t}|^{-\delta}$, for any
Hajime Sotani, Ankit Kumar
Dark matter admixed neutron stars provide a promising avenue for observationally probing the dark matter characteristic. In this study, we examine non-radial oscillations in neutron stars containing self-interacting dark matter, which interacts with normal matter exclusively via gravity. To achieve this, we derive a new set of perturbation equations for a mu
ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models
cs.CLHao Chen, Haoze Li, Zhiqing Xiao, Lirong Gao
Aligning general-purpose large language models (LLMs) to downstream tasks often incurs significant training adjustment costs. Prior research has explored various avenues to enhance alignment efficiency, primarily through minimal-data training or data-driven activations to identify key attention heads. However, these approaches inherently introduce data depen
Lexiang Hu, Yikang Li, Zhouchen Lin
The explicit governing equation is one of the simplest and most intuitive forms for characterizing physical laws. However, directly discovering partial differential equations (PDEs) from data poses significant challenges, primarily in determining relevant terms from a vast search space. Symmetry, as a crucial prior knowledge in scientific fields, has been wi
Mohammad Furquan, Tejinder P. Singh, P Samuel Wesley
The $E_8 \otimes E_8$ octonionic theory of unification suggests that our universe is six-dimensional and that the two extra dimensions are time-like. These time-like extra dimensions, in principle, offer an explanation of the quantum nonlocality puzzle, also known as the EPR paradox. Quantum systems access all six dimensions, whereas classical systems such a
Andrea Alessandrelli, Adriano Barra, Andrea Ladiana, Andrea Lepre
This paper introduces a learning framework for Three-Directional Associative Memory (TAM) models, extending the classical Hebbian paradigm to both supervised and unsupervised protocols within an hetero-associative setting. These neural networks consist of three interconnected layers of binary neurons interacting via generalized Hebbian synaptic couplings tha
Qing Li, Runze Gan, James R. Hopgood, Michael E. Davies
In this paper, we present a novel distributed expectation propagation algorithm for multiple sensors, multiple objects tracking in cluttered environments. The proposed framework enables each sensor to operate locally while collaboratively exchanging moment estimates with other sensors, thus eliminating the need to transmit all data to a central processing no
Self-patterning of Liquid Field's Metal for Enhanced Performance of Two-dimensional Semiconductor
cond-mat.mtrl-sciKwanghee Han, Heeyeon Lee, Minseong Kwon, Vinod Menon
Two-dimensional (2D) van der Waals semiconductors show promise for atomically thin flexible and transparent optoelectronic devices in future technologies.However, developing high-performance field-effect transistors (FETs) based on 2D materials is impeded by two key challenges, the high contact resistance at the 2D semiconductors-metal interface and the limi
Genie Centurion: Accelerating Scalable Real-World Robot Training with Human Rewind-and-Refine Guidance
cs.ROWenhao Wang, Jianheng Song, Chiming Liu, Jiayao Ma
While Vision-Language-Action (VLA) models show strong generalizability in various tasks, real-world deployment of robotic policy still requires large-scale, high-quality human expert demonstrations. However, data collection via human teleoperation requires continuous operator attention, which is costly, hard to scale. To address this, we propose Genie Centur
William Xie, Enora Rice, Nikolaus Correll
Humans learn how and when to apply forces in the world via a complex physiological and psychological learning process. Attempting to replicate this in vision-language models (VLMs) presents two challenges: VLMs can produce harmful behavior, which is particularly dangerous for VLM-controlled robots which interact with the world, but imposing behavioral safegu
David K. Zhang, Alex Aiken
Floating-point accumulation networks (FPANs) are key building blocks used in many floating-point algorithms, including compensated summation and double-double arithmetic. FPANs are notoriously difficult to analyze, and algorithms using FPANs are often published without rigorous correctness proofs. In fact, on at least one occasion, a published error bound fo
Exploring temporal dynamics in digital trace data: mining user-sequences for communication research
cs.SIYangliu Fan, Jakob Ohme, Lion Wedel
Communication is commonly considered a process that is dynamically situated in a temporal context. However, there remains a disconnection between such theoretical dynamicality and the non-dynamical character of communication scholars' preferred methodologies. In this paper, we argue for a new research framework that uses computational approaches to leverage
From Output to Evaluation: Does Raw Instruction-Tuned Code LLMs Output Suffice for Fill-in-the-Middle Code Generation?
cs.SEWasi Uddin Ahmad, Somshubra Majumdar, Boris Ginsburg
Post-processing is crucial for the automatic evaluation of LLMs in fill-in-the-middle (FIM) code generation due to the frequent presence of extraneous code in raw outputs. This extraneous generation suggests a lack of awareness regarding output boundaries, requiring truncation for effective evaluation. The determination of an optimal truncation strategy, how
Amir Mafi, Rando Rasul Qadir
Let $R=K[x_1,\ldots, x_n]$ be the polynomial ring in $n$ variables over a field $K$ and let $I$ be a monomial ideal of $R$. In this paper, we present an explicit formula for the Betti numbers of almost complete intersection monomial ideals, which enables a rapid construction of their minimal free resolutions. In addition, we characterize the Cohen-Macaulayne
Think Twice before Adaptation: Improving Adaptability of DeepFake Detection via Online Test-Time Adaptation
cs.CVHong-Hanh Nguyen-Le, Van-Tuan Tran, Dinh-Thuc Nguyen, Nhien-An Le-Khac
Deepfake (DF) detectors face significant challenges when deployed in real-world environments, particularly when encountering test samples deviated from training data through either postprocessing manipulations or distribution shifts. We demonstrate postprocessing techniques can completely obscure generation artifacts presented in DF samples, leading to perfo
Nazanin Mohammadi Sepahvand, Anvith Thudi, Berivan Isik, Ashmita Bhattacharyya
We present a principled, per-instance approach to quantifying the difficulty of unlearning via fine-tuning. We begin by sharpening an analysis of noisy gradient descent for unlearning (Chien et al., 2024), obtaining a better utility-unlearning tradeoff by replacing worst-case privacy loss bounds with per-instance privacy losses (Thudi et al., 2024), each of
Spinodal and Equilibrium Global Phase Diagram of the d=3 Merged Potts-Cubic-Clock Model: First-Order Equilibrium and Second-Order Spinodal Boundaries with Hidden Topologies from Renormalization-Group Theory
cond-mat.stat-mechUmut Acikel, A. Nihat Berker
A model that merges the Potts, cubic, and clock models is studied in spatial dimension d=3 by renormalization-group theory. Effective vacancies are included in the renormalization-group initial conditions. In the global phase diagram, 5 different ordered phases, namely ferromagnetic, antiferromagnetic, ferrimagnetic, antiferrimagnetic, axial, and a disordere
A physics-guided smoothing method for material modeling with digital image correlation (DIC) measurements
eess.IVJihong Wang, Chung-Hao Lee, William Richardson, Yue Yu
In this work, we present a novel approach to process the DIC measurements of multiple biaxial stretching protocols. In particular, we develop a optimization-based approach, which calculates the smoothed nodal displacements using a moving least-squares algorithm subject to positive strain constraints. As such, physically consistent displacement and strain fie
Xinbao Qiao, Ningning Ding, Yushi Cheng, Meng Zhang
Machine unlearning, as a post-hoc processing technique, has gained widespread adoption in addressing challenges like bias mitigation and robustness enhancement, colloquially, machine unlearning for fairness and robustness. However, existing non-privacy unlearning-based solutions persist in using binary data removal framework designed for privacy-driven motiv
Geometry Aware Operator Transformer as an Efficient and Accurate Neural Surrogate for PDEs on Arbitrary Domains
cs.LGShizheng Wen, Arsh Kumbhat, Levi Lingsch, Sepehr Mousavi
The very challenging task of learning solution operators of PDEs on arbitrary domains accurately and efficiently is of vital importance to engineering and industrial simulations. Despite the existence of many operator learning algorithms to approximate such PDEs, we find that accurate models are not necessarily computationally efficient and vice versa. We ad
Yahao Fan, Tianxiang Gui, Kaiyang Ji, Shutong Ding
Achieving versatile humanoid locomotion with a single policy presents a critical scalability challenge. Prevailing methods often rely on distilling multiple terrain-specific teacher policies into a unified student policy. However, while such distillation captures basic locomotion primitives, it struggles to organically compose these skills to adapt to comple
Noah Broestl, Benjamin Lange, Cristina Voinea, Geoff Keeling
Instruction-tuned Large Language Models (LLMs) are increasingly deployed as AI Assistants in firms for support in cognitive tasks. These AI assistants carry embedded perspectives which influence factors across the firm including decision-making, collaboration, and organizational culture. This paper argues that firms must align the perspectives of these AI As
Benjamin Bennetzen, Peter Buus Steffensen, Hans Hüttel, Nikolaj Rossander Kristensen
In this paper, we present a generalization of a syntax-directed editor calculus, which can be used to instantiate a specialized syntax-directed editor for any language, given by some abstract syntax. The editor calculus guarantees the absence of syntactical errors while allowing incomplete programs. The generalized editor calculus is then encoded into a simp
Kazuki Egashira, Robin Staab, Mark Vero, Jingxuan He
With the increasing size of frontier LLMs, post-training quantization has become the standard for memory-efficient deployment. Recent work has shown that basic rounding-based quantization schemes pose security risks, as they can be exploited to inject malicious behaviors into quantized models that remain hidden in full precision. However, existing attacks ca
Yiding Wang, Fauxu Meng, Xuefeng Zhang, Fan Jiang
Existing parameter-efficient fine-tuning (PEFT) methods for large language models (LLMs), such as LoRA and PiSSA, constrain model updates to low-rank subspaces, limiting their expressiveness and leading to suboptimal performance on complex tasks. To address this, we introduce High-rank Distributed PiSSA (HD-PiSSA), a distributed PEFT approach that initialize
R. S. Wilde, G. F. Gribakin, I. I. Fabrikant
We suggest that the observed large annihilation rates of ortho-positronium ($o$-Ps) in halogen gases are due to the process of dissociative Ps attachment, ${\rm Ps} + X_2 \to {\rm Ps}X + X$, where $X$ stands for a halogen atom. This process is similar to dissociative electron attachment which leads to formation of negative ions. We calculate the cross sectio
Jiayu Wang, Yang Jiao, Yue Yu, Tianwen Qian
Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for image generation. However, current benchmarks often lack the necessary breadth and depth to fully evaluate the diverse capabilities of these models. To overcome this limitation, w
From Proxies to Fields: Spatiotemporal Reconstruction of Global Radiation from Sparse Sensor Sequences
cs.LGKazuma Kobayashi, Samrendra Roy, Seid Koric, Diab Abueidda
Accurate reconstruction of latent environmental fields from sparse and indirect observations is a foundational challenge across scientific domains-from atmospheric science and geophysics to public health and aerospace safety. Traditional approaches rely on physics-based simulators or dense sensor networks, both constrained by high computational cost, latency
Mengqi Zhang, Zisheng Zhou, Xiaotian Ye, Qiang Liu
Knowledge Editing has emerged as a promising solution for efficiently updating embedded knowledge in large language models (LLMs). While existing approaches demonstrate effectiveness in integrating new knowledge and preserving the original capabilities of LLMs, they fail to maintain fine-grained irrelevant knowledge, namely facts that share the same subject
Jamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo, Matthew Jagielski
State-of-the-art membership inference attacks (MIAs) typically require training many reference models, making it difficult to scale these attacks to large pre-trained language models (LLMs). As a result, prior research has either relied on weaker attacks that avoid training references (e.g., fine-tuning attacks), or on stronger attacks applied to small model
Michal Edelstein, Hsueh-Ti Derek Liu, Mirela Ben-Chen
Learning on triangle meshes has recently proven to be instrumental to a myriad of tasks, from shape classification, to segmentation, to deformation and animation, to mention just a few. While some of these applications are tackled through neural network architectures which are tailored to the application at hand, many others use generic frameworks for triang
SPIRAL integration of generative AI in an undergraduate creative media course: effects on self-efficacy and career outcome expectations
cs.HCTroy Schotter, Saba Kawas, James Prather, Juho Leinonen
Computing education and computing students are rapidly integrating generative AI, but we know relatively little about how different pedagogical strategies for intentionally integrating generative AI affect students' self-efficacy and career interests. This study investigates a SPIRAL integration of generative AI (Skills Practiced Independently, Revisited wit
Yuedi Zhang, Shuanghao Bai, Wanqi Zhou, Zhirong Luan
Domain generalization (DG) aims to learn a model using data from one or multiple related but distinct source domains that can generalize well to unseen out-of-distribution target domains. Inspired by the success of large pre-trained vision-language models (VLMs), prompt tuning has emerged as an effective generalization strategy. However, it often struggles t
Ting-Yun Chang, Muru Zhang, Jesse Thomason, Robin Jia
Low-bit weight-only quantization significantly reduces the memory footprint of large language models (LLMs), but disproportionately affects certain examples. We analyze diverse 3-4 bit methods on LLMs ranging from 7B-70B in size and find that the quantization errors of 50 pairs of methods are strongly correlated (avg. 0.82) on FineWeb examples. Moreover, the
Nils Engler, Mathias Lindholm, Filip Lindskog, Taariq Nazar
The standard procedure to decide on the complexity of a CART regression tree is to use cross-validation with the aim of obtaining a predictor that generalises well to unseen data. The randomness in the selection of folds implies that the selected CART regression tree is not a deterministic function of the data. Moreover, the cross-validation procedure may be
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
eess.ASRui Liu, Pu Gao, Jiatian Xi, Berrak Sisman
Text-based speech editing (TSE) modifies speech using only text, eliminating re-recording. However, existing TSE methods, mainly focus on the content accuracy and acoustic consistency of synthetic speech segments, and often overlook the emotional shifts or inconsistency issues introduced by text changes. To address this issue, we propose EmoCorrector, a nove
Multi-Layer Backward Joint Model for Dynamic Prediction of Clinical Events with Multivariate Longitudinal Predictors of Mixed Types
stat.MEWenhao Li, Zhe Yin, Liang Li
Dynamic prediction of time-to-event outcomes using longitudinal data is highly useful in clinical research and practice. A common strategy is the joint modeling of longitudinal and time-to-event data. The shared random effect model has been widely studied for this purpose. However, it can be computationally challenging when applied to problems with a large n
Tianyang Li, Anping Huang, Baoyi Chen
We employ the linear response theory to calculate the polarization rate of heavy quark spin in the presence of a strong magnetic field and the hot QCD matter, both of which are simultaneously generated in relativistic heavy-ion collisions. The hot QCD medium is simplified as a fermionic system consisting of only quarks. The spin of heavy quarks can be polari
Yanjie Li, Wenxuan Zhang, Xinqi Lyu, Yihao Liu
Recently, text-to-image diffusion models have been widely used for style mimicry and personalized customization through methods such as DreamBooth and Textual Inversion. This has raised concerns about intellectual property protection and the generation of deceptive content. Recent studies, such as Glaze and Anti-DreamBooth, have proposed using adversarial no
Multiple Wasserstein Gradient Descent Algorithm for Multi-Objective Distributional Optimization
cs.LGDai Hai Nguyen, Hiroshi Mamitsuka, Atsuyoshi Nakamura
We address the optimization problem of simultaneously minimizing multiple objective functionals over a family of probability distributions. This type of Multi-Objective Distributional Optimization commonly arises in machine learning and statistics, with applications in areas such as multiple target sampling, multi-task learning, and multi-objective generativ
Zitong Wang, Hang Zhao, Qianyu Zhou, Xuequan Lu
Diffusion models have recently motivated great success in many generation tasks like object removal. Nevertheless, existing image decomposition methods struggle to disentangle semi-transparent or transparent layer occlusions due to mask prior dependencies, static object assumptions, and the lack of datasets. In this paper, we delve into a novel task: Layer-W
Yitian Yuan, Qianyue He
Recent works demonstrate the advantages of hardware rasterization for 3D Gaussian Splatting (3DGS) in forward-pass rendering through fast GPU-optimized graphics and fixed memory footprint. However, extending these benefits to backward-pass gradient computation remains challenging due to graphics pipeline constraints. We present a differentiable hardware rast
Shutong Ding, Ke Hu, Shan Zhong, Haoyang Luo
Recent advances in reinforcement learning (RL) have demonstrated the powerful exploration capabilities and multimodality of generative diffusion-based policies. While substantial progress has been made in offline RL and off-policy RL settings, integrating diffusion policies into on-policy frameworks like PPO remains underexplored. This gap is particularly si
Towards an automatic method for generating topical vocabulary test forms for specific reading passages
cs.CLMichael Flor, Zuowei Wang, Paul Deane, Tenaha O'Reilly
Background knowledge is typically needed for successful comprehension of topical and domain specific reading passages, such as in the STEM domain. However, there are few automated measures of student knowledge that can be readily deployed and scored in time to make predictions on whether a given student will likely be able to understand a specific content ar
How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
cs.CLMinglai Yang, Ethan Huang, Liang Zhang, Mihai Surdeanu
We introduce Grade School Math with Distracting Context (GSM-DC), a synthetic benchmark to evaluate Large Language Models' (LLMs) reasoning robustness against systematically controlled irrelevant context (IC). GSM-DC constructs symbolic reasoning graphs with precise distractor injections, enabling rigorous, reproducible evaluation. Our experiments demonstrat
Ruichen Zhang, Rana Muhammad Shahroz Khan, Zhen Tan, Dawei Li
Data-centric distillation, including data augmentation, selection, and mixing, offers a promising path to creating smaller, more efficient student Large Language Models (LLMs) that retain strong reasoning abilities. However, there still lacks a comprehensive benchmark to systematically assess the effect of each distillation approach. This paper introduces DC
Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding
cs.LGAlexander Conzelmann, Robert Bamler
The ever-growing size of neural networks poses serious challenges on resource-constrained devices, such as embedded sensors. Compression algorithms that reduce their size can mitigate these problems, provided that model performance stays close to the original. We propose a novel post-training compression framework that combines rate-aware quantization with e
Duo Li, Zuhao Yang, Xiaoqin Zhang, Ling Shao
Visual token pruning aims to compress and prune redundant visual tokens which play a critical role in efficient inference with large vision-language models (LVLMs). However, most existing work estimates visual redundancy using a single metric, such as cross-modal attention or visual token similarity. We show that visual token diversity and task-specific toke
Accelerated Bayesian calibration and uncertainty quantification of RANS turbulence model parameters for stratified atmospheric boundary layer flows
physics.flu-dynE. Y. Shin, M. F. Howland
In operational weather models, the effects of turbulence in the atmospheric boundary layer (ABL) on the resolved flow are modeled using turbulence parameterizations. These parameterizations typically use a predetermined set of model parameters that are tuned to limited data from canonical flows. Using these fixed parameters results in deterministic predictio
Xiaolu Chen, Chenghao Huang, Yanru Zhang, Hao Wang
With the proliferation of smart grids, smart cities face growing challenges due to cyber-attacks and sophisticated electricity theft behaviors, particularly in residential photovoltaic (PV) generation systems. Traditional Electricity Theft Detection (ETD) methods often struggle to capture complex temporal dependencies and integrating multi-source data, limit
Laurance Fakih, Andrei Halanay
Antimicrobial Resistance (RAM) poses a significant threat to global public health, making important medicines less useful. While the medical and biological reasons behind RAM are well studied, we still don't know enough about how false health information affects people's actions, which can speed up RAM. This study presents a new mathematical model to investi
Few-Shot Optimization for Sensor Data Using Large Language Models: A Case Study on Fatigue Detection
cs.CLElsen Ronando, Sozo Inoue
In this paper, we propose a novel few-shot optimization with HED-LM (Hybrid Euclidean Distance with Large Language Models) to improve example selection for sensor-based classification tasks. While few-shot prompting enables efficient inference with limited labeled data, its performance largely depends on the quality of selected examples. HED-LM addresses thi
The Dual Horizon: A Rendezvous of Computing and Communication Services at the Optical Layer in Optical Computing-Communication Integrated Network
cs.NIDao Thanh Hai, Isaac Woungang
With the significant advancements in optical computing platforms recently capable of performing various primitive operations, a seamless integration of optical computing into very fabric of optical communication links is envisioned, paving the way for the advent of \textit{optical computing-communication integrated network}, which provides computing services
Haolin Yang, Hakaze Cho, Yiqiao Zhong, Naoya Inoue
The unusual properties of in-context learning (ICL) have prompted investigations into the internal mechanisms of large language models. Prior work typically focuses on either special attention heads or task vectors at specific layers, but lacks a unified framework linking these components to the evolution of hidden states across layers that ultimately produc
Haiqi Wu, Kai Xu
This paper investigates the relationship between categorical entropy and von Neumann entropy of quantum lattices. We begin by studying the von Neumann entropy, proving that the average von Neumann entropy per site converges to the logarithm of an algebraic integer in the low-temperature and thermodynamic limits. Next, we turn to categorical entropy. Given an
Agent-Based Decentralized Energy Management of EV Charging Station with Solar Photovoltaics via Multi-Agent Reinforcement Learning
eess.SYJiarong Fan, Chenghao Huang, Hao Wang
In the pursuit of energy net zero within smart cities, transportation electrification plays a pivotal role. The adoption of Electric Vehicles (EVs) keeps increasing, making energy management of EV charging stations critically important. While previous studies have managed to reduce energy cost of EV charging while maintaining grid stability, they often overl
Ubiquity of rotational symmetry breaking in superconducting films, from Fe(Te,Se)/Bi$_2$Te$_3$ to Nb, and the effect of measurement geometry
cond-mat.supr-conDebarghya Mallick, Hee Taek Yi, Xiaoyu Yuan, Seongshik Oh
FeTe$_{0.5}$Se$_{0.5}$/Bi$_2$Te$_3$ heterostructure is a promising new platform in the journey toward topological quantum computation, considering that first, FeTe$_{0.5}$Se$_{0.5}$ is itself known to be a topological superconductor (TSC) and second, the heterostructure has topological interface states that can be proximitized into TSC even if FTS fails to b
Clustering analysis of BOSS-CMASS galaxies with semi-analytical model for galaxy formation and halo occupation distribution
astro-ph.COZhongxu Zhai, Andrew Benson, Yun Wang
The spatial distribution of massive and luminous galaxies have provided important constraints on the fundamental cosmological parameters and physical processes governing galaxy formation. In this work, we construct and compare independent galaxy-halo connection models in the application of clustering measurement at non-linear scales of BOSS-CMASS galaxies. I
Season-Independent PV Disaggregation Using Multi-Scale Net Load Temporal Feature Extraction and Weather Factor Fusion
eess.SPXiaolu Chen, Chenghao Huang, Yanru Zhang, Hao Wang
With the advancement of energy Internet and energy system integration, the increasing adoption of distributed photovoltaic (PV) systems presents new challenges on smart monitoring and measurement for utility companies, particularly in separating PV generation from net electricity load. Existing methods struggle with feature extraction from net load and captu
Peijie Yu, Yifan Yang, Jinjian Li, Zelong Zhang
Agents based on large language models leverage tools to modify environments, revolutionizing how AI interacts with the physical world. Unlike traditional NLP tasks that rely solely on historical dialogue for responses, these agents must consider more complex factors, such as inter-tool relationships, environmental feedback and previous decisions, when making
Umar Marikkar, Syed Sameed Husain, Muhammad Awais, Sara Atito
Immunohistochemical (IHC) images reveal detailed information about structures and functions at the subcellular level. However, unlike natural images, IHC datasets pose challenges for deep learning models due to their inconsistencies in channel count and configuration, stemming from varying staining protocols across laboratories and studies. Existing approach
Adoubi Vincent De Paul Adombi
Scientific machine learning (SciML) provides a structured approach to integrating physical knowledge into data-driven modeling, offering significant potential for advancing hydrological research. In recent years, multiple methodological families have emerged, including physics-informed machine learning, physics-guided machine learning, hybrid physics-machine
Resistive Collapse of 2D Non-rotating Magnetized Isothermal Toroids: Formation of Pseudodisks
astro-ph.SRYa-Chi Wang, Hsien Shang, Ruben Krasnopolsky
The collapse of singular magnetized toroids (Li & Shu 1996) is a natural representation of an early phase in star formation, bridging the prestellar and protostellar phases of the collapse of molecular cloud cores. We revisit the collapse study of Allen et al. (2003b), now with explicit nonideal MHD (Ohmic diffusivity $\eta$) and higher resolution using a co
Han Li, Hu Han, S. Kevin Zhou
Natural images exhibit label diversity (clean vs. noisy) in noisy-labeled image classification and prevalence diversity (abundant vs. sparse) in long-tailed image classification. Similarly, medical images in universal lesion detection (ULD) exhibit substantial variations in image quality, encompassing attributes such as clarity and label correctness. How to
Greg Bodwin, Tuong Le
When regularity lemmas were first developed in the 1970s, they were described as results that promise a partition of any graph into a ``small'' number of parts, such that the graph looks ``similar'' to a random graph on its edge subsets going between parts. Regularity lemmas have been repeatedly refined and reinterpreted in the years since, and the modern pe
Kenza Abela, Shaima Abidrabbu, Ayoub Ammar Boudjelal, Huseyin Arslan
Two critical approaches have emerged in the literature for the successful realization of 6G wireless networks: the coexistence of multiple waveforms and the adoption of non-orthogonal multiple access. These strategies hold transformative potential for addressing the limitations of current systems and enabling the robust and scalable design of next-generation
Haonan Dong, Wenhao Zhu, Guojie Song, Liang Wang
Low-Rank Adaptation (LoRA) is a widely adopted parameter-efficient fine-tuning (PEFT) method validated across NLP and CV domains. However, LoRA faces an inherent low-rank bottleneck: narrowing its performance gap with full finetuning requires increasing the rank of its parameter matrix, resulting in significant parameter overhead. Recent linear LoRA variants
Rahul Barthwal, Christian Rohde, Yue Wang
In this paper, a second-order generalized Riemann problem (GRP) solver is developed for a two-layer thin film model. Extending the first-order Godunov approach, the solver is used to construct a temporal-spatial coupled second-order GRP-based finite-volume method. Numerical experiments including comparisons to MUSCL finite-volume schemes with Runge-Kutta tim
Junyong Kang, Seohyun Lim, Kyungjune Baek, Hyunjung Shim
Aligning text-to-image (T2I) diffusion models with human preferences has emerged as a critical research challenge. While recent advances in this area have extended preference optimization techniques from large language models (LLMs) to the diffusion setting, they often struggle with limited exploration. In this work, we propose a novel and orthogonal approac