March 2026 arXiv papers — page 20
Showing 1,901–2,000 of 25,974 papers
RCLRec: Reverse Curriculum Learning for Modeling Sparse Conversions in Generative Recommendation
cs.IRYulei Huang, Hao Deng, Haibo Xing, Jinxin Hu
Conversion objectives in large-scale recommender systems are sparse, making them difficult to optimize. Generative recommendation (GR) partially alleviates data sparsity by organizing multi-type behaviors into a unified token sequence with shared representations, but conversion signals remain insufficiently modeled. While recent behavior-aware GR models enco
Parham Pourdavood
Constitutional AI (CAI) aligns language models with explicitly stated normative principles, offering a transparent alternative to implicit alignment through human feedback alone. However, because constitutions are authored by specific groups of people, the resulting models may reflect particular cultural perspectives. We investigate this question by evaluati
Q-DIVER: Integrated Quantum Transfer Learning and Differentiable Quantum Architecture Search with EEG Data
quant-phJunghoon Justin Park, Yeonghyeon Park, Jiook Cha
Integrating quantum circuits into deep learning pipelines remains challenging due to heuristic design limitations. We propose Q-DIVER, a hybrid framework combining a large-scale pretrained EEG encoder (DIVER-1) with a differentiable quantum classifier. Unlike fixed-ansatz approaches, we employ Differentiable Quantum Architecture Search to autonomously discov
Joint Time-Phase Synchronization for Distributed Sensing Networks via Feature-Level Hyper-Plane Regression
eess.SPKailun Tian, Kaili Jiang, Dechang Wang, Yuxin Zhao
Achieving coherent integration in distributed Internet of Things (IoT) sensing networks requires precise synchronization to jointly compensate clock offsets and radio-frequency (RF) phase errors. Conventional two-step protocols suffer from time-phase coupling, where residual timing offsets degrade phase coherence. This paper proposes a generalized hyper-plan
MedLoc-R1: Performance-Aware Curriculum Reward Scheduling for GRPO-Based Medical Visual Grounding
cs.CVGuangjing Yang, Ziyuan Qin, Chaoran Zhang, Chenlin Du
Medical visual grounding serves as a crucial foundation for fine-grained multimodal reasoning and interpretable clinical decision support. Despite recent advances in reinforcement learning (RL) for grounding tasks, existing approaches such as Group Relative Policy Optimization~(GRPO) suffer from severe reward sparsity when directly applied to medical images,
Haoxiang Jia, Earl T. Barr, Sergey Mechtaev
Large Language Models (LLMs) are now capable of resolving real-world GitHub issues. However, current approaches overapproximate the code context and suffer from two compounding problems: the prohibitive cost of processing massive inputs, and low effectiveness as noise floods the context window and distracts the model from the bug-fixing signal. Existing comp
Sofia Brenner, Jiří Fink
We present an algorithm that enumerates all ideals of an input poset with constant delay in Gray code order, i.e., such that consecutively visited ideals differ in at most three elements. This answers a long-standing open problem posed by Pruesse and Ruskey, and improves upon previous algorithms by Pruesse and Ruskey, Squire, Habib, Medina, Nourine and Stein
Shoujin Wang, Mingze Ni, Wei Liu, Victor W. Chu
Livestock growth prediction is essential for optimising farm management and improving the efficiency and sustainability of livestock production, yet it remains underexplored due to limited large-scale datasets and privacy concerns surrounding farm-level data. Existing biophysical models rely on fixed formulations, while most machine learning approaches are t
$AutoDrive\text{-}P^3$: Unified Chain of Perception-Prediction-Planning Thought via Reinforcement Fine-Tuning
cs.ROYuqi Ye, Zijian Zhang, Junhong Lin, Shangkun Sun
Vision-language models (VLMs) are increasingly being adopted for end-to-end autonomous driving systems due to their exceptional performance in handling long-tail scenarios. However, current VLM-based approaches suffer from two major limitations: 1) Some VLMs directly output planning results without chain-of-thought (CoT) reasoning, bypassing crucial percepti
Graph Vector Field: A Unified Framework for Multimodal Health Risk Assessment from Heterogeneous Wearable and Environmental Data Streams
cs.LGSilvano Coletti, Francesca Fallucchi
Digital health research has advanced dynamic graph-based disease models, topological learning on simplicial complexes, and multimodal mixture-of-experts architectures, but these strands remain largely disconnected. We propose Graph Vector Field (GVF), a framework that models health risk as a vector-valued field on time-varying simplicial complexes, coupling
Seunghun Oh, Unsang Park
Cross-attention is the primary interface through which text conditions latent diffusion models, yet its step-wise multi-resolution dynamics remain under-characterized, limiting principled training-free control. We cast diffusion cross-attention as a spatiotemporal signal on the latent grid by summarizing token-softmax weights into token-agnostic concentratio
Simon Kuang, Yuezhu Xu, S. Sivaranjani, Xinfan Lin
The global Lipschitz constant of a neural network is related to robustness and generalization, yet unlike in many classical models, it is not plainly legible from the parameters. This has motivated sophisticated verification algorithms, especially semidefinite programming (SDP) based on incremental quadratic constraints on the activation functions, to improv
Yuichi Goto, Gaspard Bernard
Recently, several spectra have emerged, designed to encapsulate the distributional characteristics of non-Gaussian stationary processes. This article introduces parametric families of generalized spectra based on the characteristic function, alongside inference procedures enabling $\sqrt{n}$-consistent estimation of the unknown parameters in a broad class of
Keiho Matsumoto
We study the question of whether the vanishing of additive invariants characterizes phantomness for smooth proper dg categories admitting geometric realizations. More precisely, let $X$ be a smooth proper variety over a field $k$, and let $\sT\subset \perfdg(X)$ be a $k$-linear admissible full dg subcategory. We construct a non-compact motive $\sM(\sT)\in \D
Contour-Guided Query-Based Feature Fusion for Boundary-Aware and Generalizable Cardiac Ultrasound Segmentation
cs.CVZahid Ullah, Sieun Choi, Jihie Kim
Accurate cardiac ultrasound segmentation is essential for reliable assessment of ventricular function in intelligent healthcare systems. However, echocardiographic images are challenging due to low contrast, speckle noise, irregular boundaries, and domain shifts across devices and patient populations. Existing methods, largely based on appearance-driven lear
Finite-blocklength performance of polar wiretap codes under a total variation secrecy constraint
cs.ITLaura Luzzi, Valerio Bioglio
We study the performance of polarizing codes over a degraded symmetric wiretap channel under a total variation distance (TVD) secrecy constraint. We show that the leakage can be bounded by the sum of the TVDs of the bit-channels corresponding to the confidential and frozen bits. In the asymptotic regime, this gives a new criterion to design wiretap codes wit
Leonardo Bassanini, Ludovico Biancardi, Alfio Ferrara, Andrea Gamberini
The digitisation of historical documents has traditionally been conceived as a process limited to character-level transcription, producing flat text that lacks the structural and semantic information necessary for substantive computational analysis. We present VERITAS (Vision-Enhanced Reading, Interpretation, and Transcription of Archival Sources), a modular
Shape, regolith size and thickness, SMFe^0 content, and spectral type of Tianwen-2 target asteroid (469219) Kamo'oalewa
astro-ph.EPPengfei Zhang, Guozheng Zhang, Yongxiong Zhang, Marco Fenucci
China's Tianwen-2 spacecraft will return samples from the near-Earth asteroid (469219) Kamo'oalewa. We previously reported that Kamo'oalewa develops an LL-chondrite-compositional, highly space-weathered surface. This study aims to estimate Kamo'oalewa's shape, regolith grain size and thickness, sub-micrometer iron (SMFe0) content, and spectral type. Using th
InconLens: Interactive Visual Diagnosis of Behavioral Inconsistencies in LLM-based Agentic Systems
cs.HCShuo Yan, Xiaolin Wen, Shaolun Ruan, Yanjie Zhang
Large Language Model (LLM)-based agentic systems have shown growing promise in tackling complex, multi-step tasks through autonomous planning, reasoning, and interaction with external environments. However, the stochastic nature of LLM generation introduces intrinsic behavioral inconsistency: the same agent may succeed in one execution but fail in another un
Chunhang Zheng, Tongda Xu, Mingli Xie, Yan Wang
Raw images preserve linear sensor measurements and high bit-depth information crucial for advanced vision tasks and photography applications, yet their storage remains challenging due to large file sizes, varying bit depths, and sensor-dependent characteristics. Existing learned lossless compression methods mainly target 8-bit sRGB images, while raw reconstr
Daniel A. Goldston, Ade Irma Suriajaya
In 1973 Montgomery proved, assuming the Riemann Hypothesis (RH), that asymptotically at least 2/3 of zeros of the Riemann zeta-function are simple zeros. In a previous note (arXiv:2511.20059 [math.NT]) we showed how RH can be replaced with a general estimate for a double sum over zeros, and this allows one to then obtain results on zeros that are both simple
Tianwen-2 target asteroid (469219) Kamo'oalewa probably develops an Itokawa-compositional but ultra-highly space-weathered surface
astro-ph.EPPengfei Zhang, Guozheng Zhang, Zichen Wei, Mikael Granvik
China's Tianwen-2 mission plans to return samples from a small, rapidly spinning Earth quasi-satellite (469219) Kamo'oalewa. Previous studies linked Kamo'oalewa to lunar composition and origin. Here, we propose another scenario. We reanalyzed the reflectance spectrum of Kamo'oalewa and obtained an absorption band center at 1.001+-0.028 um (error is 1sigma),
Zili Zhang, Yinmin Zhong, Chengxu Yang, Chao Jin
Agentic Reinforcement Learning (RL) enables LLMs to solve complex tasks by alternating between a data-collection rollout phase and a policy training phase. During rollout, the agent generates trajectories, i.e., multi-step interactions between LLMs and external tools. Yet, frequent tool calls induce long-tailed trajectory generation that bottlenecks rollouts
Kacper Kluk, Hung Le, Wojciech Nadara, Marcin Pilipczuk
A furthest neighbor data structure on a metric space $(V,\mathrm{dist})$ and a set $P \subseteq V$ answers the following query: given $v \in V$, output $p \in P$ maximizing $\mathrm{dist}(v,p)$; in the approximate version, it is allowed to report any $p \in P$ with $\mathrm{dist}(v,p) \geq (1-\varepsilon)\max_{p' \in P} \mathrm{dist}(v,p')$ for an accuracy p
Eigenvalue-based Linear Stability Analysis of Intrinsic Instabilities in Laminar Flames
physics.flu-dynThomas Ludwig Kaiser, Peter Munch, Sandra May, Thorsten Zirwes
Intrinsic instabilities of laminar premixed flames play an important role in the dynamics of hydrogen combustion and in the development of predictive models for reacting flows. However, determining their dispersion relations typically relies either on simplified analytical descriptions of the flame front or on computationally expensive direct numerical simul
Developing and characterizing a new-generation regolith simulant "IGCAS-AST01" for the Tianwen-2 target asteroid (469219) Kamo'oalewa
astro-ph.EPPengfei Zhang, Zichen Wei, Takahiro Hiroi, Jin Zhao
China plans to return samples from the near-Earth asteroid (469219) Kamo'oalewa, which we previously identified as an LL-chondrite-compositional, highly space-weathered object with fine-grained regolith. In this study, we developed 10 mL of Kamo'oalewa regolith simulant, designated "IGCAS-AST01", by irradiating LL5/6 chondrite (Kheneg Ljou^ad) powder with a
Kaiyu Zheng, Wei Gao, Huiming Zheng
Octree-based context learning has recently become a leading method in point cloud compression. However, its potential on lossy compression remains undiscovered. The traditional lossy compression paradigm using lossless octree representation with quantization step adjustment may result in severe distortions due to massive missing points in quantization. There
Mark D. Gould, Artem Pulemotov, Jorgen Rasmussen, Yang Zhang
We classify all irreducible highest-weight unitary modules over the non-compact real form $\mathfrak{u}(p,q|n)$ of the general linear Lie superalgebra $\mathfrak{gl}_{p+q|n}$. The classification is given by explicit necessary and sufficient conditions on the highest weights, and our approach combines the Howe duality for $\mathfrak{gl}_{p+q|n}$ with a quadra
He Yang, Dongyi Lv, Song Ma, Wei Xi
Dataset Condensation (DC) is a data-efficient learning paradigm that synthesizes small yet informative datasets, enabling models to match the performance of full-data training. However, recent work exposes a critical vulnerability of DC to backdoor attacks, where malicious patterns (\textit{e.g.}, triggers) are implanted into the condensation dataset, induci
Alexander Prutsch, Christian Fruhwirth-Reisinger, David Schinagl, Horst Possegger
In dynamic traffic environments, motion forecasting models must be able to accurately estimate future trajectories continuously. Streaming-based methods are a promising solution, but despite recent advances, their performance often degrades when exposed to heterogeneous observation lengths. To address this, we propose a novel streaming-based motion forecasti
Seo-Won Chang, Myungshin Im, Mankeun Jeong, Joonho Kim
We present the first public data release (DR1) of the KMTNet Synoptic Survey of Southern Sky (KS4). This deep, wide-field imaging survey covers a southern footprint of -85$^{\circ}$ < Decl. < -28.8$^{\circ}$ in the $B$, $V$, $R$, and $I$ bands using a network of three 1.6-m telescopes. Although primarily designed to secure reference imaging for gravitational
Zefeng He, Siyuan Huang, Xiaoye Qu, Yafu Li
Recent multimodal generation models have achieved remarkable progress on general-purpose generation tasks, yet continue to struggle with complex instructions and specialized downstream tasks. Inspired by the success of advanced agent frameworks such as Claude Code, we propose \textbf{GEMS} (Agent-Native Multimodal \textbf{GE}neration with \textbf{M}emory and
Souvik Mandal, Ankur Sarkar
Recently, the Mac\'ias topology has been generalized over integral domains that are not fields, to furnish a topological proof of the infinitude of prime elements under the assumption that the set of units is finite or not open. In this article, we remove this cardinality assumption completely by using the Jacobson radical. We prove that in any semiprimitive
Kexin Huang, Liwei Fan, Botian Jiang, Yaozhou Jiang
Voice design from natural language aims to generate speaker timbres directly from free-form textual descriptions, allowing users to create voices tailored to specific roles, personalities, and emotions. Such controllable voice creation benefits a wide range of downstream applications-including storytelling, game dubbing, role-play agents, and conversational
Andreas Bluhm, Gereon Koßmann, René Schwonnek
Device-independent quantum key distribution (DIQKD) provides a model of quantum key distribution with minimal assumptions and highly abstract theoretical building blocks. Although DIQKD frees us from detailed discussions of specific device models and associated error parameters, it replaces them with fundamental assumptions about the validity of quantum expe
Changjian Su, Yang Yang
We study the equivariant homology of the generalized Steinberg variety of type C and show that there exists a surjective algebra homomorphism from the twisted Yangian of type $\AIII_{2n}^{(\tau)}$ to it.
Junzhe Song, Ruisi He, Mi Yang, Zhengyu Zhang
Site-specific channel inference plays a critical role in the design and evaluation of next-generation wireless communication systems by considering the surrounding propagation environment. However, traditional methods are unscalable. Recently, satellite imagery has emerged as a valuable modality containing rich propagation information for AI-based channel pr
Chutian Meng, Fan Ma, Chi Zhang, Jiaxu Miao
Generating coherent and communicative visual sequences, such as image sequences and videos, remains a significant challenge for current multimodal systems. Despite advances in visual quality and the integration of world knowledge, existing models still struggle to maintain logical flow, often resulting in disjointed actions, fragmented narratives, and unclea
Roman Smirnov
Modern large language models (LLMs) are used in many business applications in general, and specifically in web search systems and applications that generate overviews of search results - LLM Overview systems. Such systems are using an LLM to select most relevant sources from search results and generate an answer to the user's query. It is known from many stu
Transformer-Based Prognostics: Enhancing Network Availability by Improved Monitoring of Optical Fiber Amplifiers
eess.SPDominic Schneider, Lutz Rapp, Christoph Ament
We enhance optical network availability and reliability through a lightweight transformer model that predicts optical fiber amplifier lifetime from condition-based monitoring data, enabling real-time, edge-level predictive maintenance and advancing deployable AI for autonomous network operation.
Control Without Control: Defining Implicit Interaction Paradigms for Autonomous Assistive Robots
cs.ROJanavi Gupta, Kavya Puthuveetil, Dimitra Tsakona, Akhil Padmanabha
Assistive robotic systems have shown growing potential to improve the quality of life of those with disabilities. As researchers explore the automation of various caregiving tasks, considerations for how the technology can still preserve the user's sense of control become paramount to ensuring that robotic systems are aligned with fundamental user needs and
Yoshikazu Giga, Ayato Kubo, Hirotoshi Kuroda, Koya Sakakibara
We consider a Kobayashi-Warren-Carter (KWC) type total variation energy with a fidelity term. Since the energy is non-convex, the profiles of minimizers are quite different from those of the original Rudin-Osher-Fatemi energy. In one-dimensional setting, we prove that KWC type energy (and its generalization) with fidelity must have a piecewise constant minim
Qin-Ru Cheng, Ke-Xiong Yan, Yuan Qiu, Yi-Tong Shi
We propose an efficient and robust protocol for the generation of entanglement between a superconducting qubit and a squeezed cavity. By applying a parametric drive to the cavity coupled to the qubit, the dynamical evolution of the system is precisely described by an anisotropic Rabi model within a squeezed reference frame. Utilizing high-order time-averagin
Wen Ting Hsieh, Alev Orfi, Dries Sels
Quantum-enhanced Markov chain Monte Carlo, a hybrid quantum-classical algorithm in which configurations are proposed by a quantum proposer and accepted or rejected by a classical algorithm, has been introduced as a possible method for robust quantum speedup. Previous work has identified competing factors that limit the algorithm's performance: the quantum dy
Melissa Yactayo, A. Pezo, J. L. Ampuero, M. Tian
Emerging orbitronics assumes long-range orbital current transport, analogous to spin currents. However, recent theory and experiments challenge this view, showing rather local characters for orbital polarization and orbit-spin conversions. We study angular momentum generated by ferromagnetic resonance and thermal gradients in Ni/(Pt)Ti/Au heterostructures. T
Koopman-based surrogate modeling for reinforcement-learning-control of Rayleigh-Benard convection
cs.LGTim Plotzki, Sebastian Peitz
Training reinforcement learning (RL) agents to control fluid dynamics systems is computationally expensive due to the high cost of direct numerical simulations (DNS) of the governing equations. Surrogate models offer a promising alternative by approximating the dynamics at a fraction of the computational cost, but their feasibility as training environments f
SIMR-NO: A Spectrally-Informed Multi-Resolution Neural Operator for Turbulent Flow Super-Resolution
cs.LGMuhammad Abid, Omer San
Reconstructing high-resolution turbulent flow fields from severely under-resolved observations is a fundamental inverse problem in computational fluid dynamics and scientific machine learning. Classical interpolation methods fail to recover missing fine-scale structures, while existing deep learning approaches rely on convolutional architectures that lack th
Muhittin Evren Aydin, Esra Dilmen, Busra Karakaya
In this paper, we introduce the notion of a prescribed angle curve in a Riemannian manifold associated with a pair $(\mathcal{V},\theta)$, where $\mathcal{V}$ is a unit vector field along the curve and $\theta$ denotes the angle between $\mathcal{V}$ and the principal normal vector of the curve. When $\mathcal{V}$ is a torse-forming vector field, we establis
Dynamical diffraction formalism for imaging time-dependent diffuse scattering from coherent phonons with Dark-Field X-ray Microscopy
cond-mat.mes-hallDarshan Chalise, Brinthan Kanesalingam, Dorian P. Luccioni, Daniel Schick
Coherent acoustic phonons, whose damping sets the upper bound of quality factors in acoustic resonators, play a critical role in advanced telecommunication and quantum information technologies. Yet, probing their decay in the GHz regime remains challenging using conventional surface-based techniques. Dark-field X-ray microscopy (DFXM) offers a solution by en
Tsukasa Iwabuchi, Hideo Kozono
We consider the 3D incompressible Euler equations in bounded domains $\Omega$ with smooth boundary $\partial\Omega$. Based on the paper by Iwabuchi, Matsuyama and Taniguchi (2019), we define the Besov space $B^s_{p, q}(A)$ by means of the Stokes operator $A$ with the Neumann boundary condition on $\partial\Omega$, and prove unique local existence theorem of
Christopher Clark, Yue Yang, Jae Sung Park, Zixian Ma
Grounding has become a fundamental capability of vision-language models (VLMs). Most existing VLMs point by generating coordinates as part of their text output, which requires learning a complicated coordinate system and results in a high token count. Instead, we propose a more intuitive pointing mechanism that directly selects the visual tokens that contain
Zhaohe Liao, Kaixun Jiang, Zhihang Liu, Yujie Wei
Although image generation has boosted various applications via its rapid evolution, whether the state-of-the-art models are able to produce ready-to-use academic illustrations for papers is still largely unexplored. Directly comparing or evaluating the illustration with VLM is native but requires oracle multi-modal understanding ability, which is unreliable
Huanxing Chen, Aditesh Kumar
Generative agent simulations operate at two scales: individual personas for character interaction, and population models for collective behavior analysis and intervention testing. We propose a third scale: meso-level simulation - interaction with group-level representations that retain grounding in rich individual experience. To enable this, we present Synon
KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness
cs.CLJinyoung Kim, Hyeongsoo Lim, Eunseo Seo, Minho Jang
Recent advances in large audio language models (LALMs) have enabled multilingual speech understanding. However, benchmarks for evaluating LALMs remain scarce for non-English languages, with Korean being one such underexplored case. In this paper, we introduce KoALa-Bench, a comprehensive benchmark for evaluating Korean speech understanding and speech faithfu
Sergio Muñiz Subiñas, Alejandro Mata Ali, Jorge Martínez Martín, Miguel Franco Hernando
This work presents a novel tensor network algorithm for solving Quadratic Unconstrained Binary Optimization (QUBO) problems, Quadratic Unconstrained Discrete Optimization (QUDO) problems, and Tensor Quadratic Unconstrained Discrete Optimization (T-QUDO) problems. The proposed algorithm is based on the MeLoCoToN methodology, which solves combinatorial optimiz
Renjie Wu, Hongdong Li, Jose M. Alvarez, Miaomiao Liu
This paper addresses the problem of dynamic scene surface reconstruction using Gaussian Splatting (GS), aiming to recover temporally consistent geometry. While existing GS-based dynamic surface reconstruction methods can yield superior reconstruction, they are typically limited to either a single object or objects with only small deformations, struggling to
Jiacheng Wang, Jinbin Huang
We prove that under five minimal axioms -- multi-dimensional quality, finite evaluation, effective optimization, resource finiteness, and combinatorial interaction -- any optimized AI agent will systematically under-invest effort in quality dimensions not covered by its evaluation system. This result establishes reward hacking as a structural equilibrium, no
Yuang Wei, Ruijia Li, Bo Jiang
While Large Language Models (LLMs) have demonstrated remarkable fluency in educational dialogues, most generative tutors primarily operate through intuitive, single-pass generation. This reliance on fast thinking precludes a dedicated reasoning workspace, forcing multiple diagnostic and strategic signals to be processed in a conflated manner. As a result, le
Vipul Arora, Arnab Bhattacharyya, Philips George John, Sayantan Sen
Over the last three decades, function testing has been extensively studied over Boolean, finite fields, and discrete settings. However, to encode the real-world applications more succinctly, function testing over the reals (where the domain and range, both are reals) is of prime importance. Recently, there have been some works in the direction of testing for
A Classification of Heterogeneity in Uncrewed Vehicle Swarms and the Effects of Its Inclusion on Overall Swarm Resilience
cs.ROAbhishek Joshi, Abhishek Phadke, Tianxing Chu, F. Antonio Medrano
Combining different types of agents in uncrewed vehicle (UV) swarms has emerged as an approach to enhance mission resilience and operational capabilities across a wide range of applications. This study offers a systematic framework for grouping different types of swarms based on three main factors: agent nature (behavior and function), hardware structure (ph
DAInfer+: Neurosymbolic Inference of API Specifications from Documentation via Embedding Models
cs.SEMaryam Masoudian, Anshunkang Zhou, Chengpeng Wang, Charles Zhang
Modern software systems heavily rely on various libraries, which require understanding the API semantics in static analysis. However, summarizing API semantics remains challenging due to complex implementations or unavailable library code. This paper presents DAInfer+, a novel approach for inferring API specifications from library documentation. We employ Na
David Cheban
The aim of this paper is to study the remotely almost periodic motions of dynamical systems and solutions of nonlinear differential equations. We establish some properties of remotely almost periodic motions and generalize the well known Amerio's theorem for abstract remotely almost periodic dynamical systems. Application of our general results for different
Is One-Shot In-Context Learning Helpful for Data Selection in Task-Specific Fine-Tuning of Multimodal LLMs?
cs.MMXiao An, Jiaxing Sun, Ting Hu, Wei He
Injecting world knowledge into pretrained multimodal large language models (MLLMs) is essential for domain-specific applications. Task-specific fine-tuning achieves this by tailoring MLLMs to high-quality in-domain data but encounters scalability challenges as datasets grow, necessitating a trade-off between performance and computational overhead. Existing d
Pulock Das, Al Amin, Kamrul Hasan, Rohan Thompson
Deep learning (DL) models have achieved strong performance in an intelligence healthcare setting, yet most existing approaches operate as black boxes and ignore the physical processes that govern tumor growth, limiting interpretability, robustness, and clinical trust. To address this limitation, we propose PhysNet, a physics-embedded DL framework that integr
Ferromagnetic resonance modulation in topological materials with bulk--boundary coexistence
cond-mat.mes-hallShun Muto, Yuya Ominato, Takeo Kato, Mamoru Matsuo
We extend ferromagnetic resonance (FMR) modulation theory to describe systems in which bulk and boundary states of topological materials coexist, with both appearing at the same energy. As an application of the formulation, we investigate the enhancement of the Gilbert damping constant on the $(110)$ surface of a $d$-wave superconductor where nodal quasipart
Sonae Hadama
In this paper, we study a class of one-dimensional nonlocal nonlinear Schr\"odinger equations on the line with nonlinearity given by a Fourier multiplier whose symbol has subcritical high-frequency growth. In terms of symbol order, this class is intermediate between the cubic nonlinear Schr\"odinger equation and the Calogero--Moser derivative nonlinear Schrd
Udita Ghosh, Dripta S. Raychaudhuri, Jiachen Li, Konstantinos Karydis
Preference-based reinforcement learning can learn effective reward functions from comparisons, but its scalability is constrained by the high cost of oracle feedback. Lightweight vision-language embedding (VLE) models provide a cheaper alternative, but their noisy outputs limit their effectiveness as standalone reward generators. To address this challenge, w
Yoonho Lee, Roshen Nair, Qizheng Zhang, Kangwook Lee
The performance of large language model (LLM) systems depends not only on model weights, but also on their harness: the code that determines what information to store, retrieve, and present to the model. Yet harnesses are still designed largely by hand, and existing text optimizers are poorly matched to this setting because they compress feedback too aggress
A domain hemivariational inequality for 2D and 3D convective Brinkman-Forchheimer extended Darcy equations
math.APJyoti Jindal, Sagar Gautam, Manil T. Mohan
This paper investigates domain hemivariational inequality problems arising from the non-stationary two- and three-dimensional convective Brinkman-Forchheimer extended Darcy (CBFeD) equations, which describe the flow of viscous incompressible fluids through saturated porous media in bounded domains. These equations may be regarded as generalized Navier-Stokes
Liang Sun
Based on the Distributed Convolutional Neural Network(DisCNN), a straightforward object detection method is proposed. The modules of the output vector of a DisCNN with respect to a specific positive class are positively monotonic with the presence probabilities of the positive features. So, by identifying all high-scoring patches across all possible scales,
Yuma Yamaoka, Seiichi Uchida, Shoji Toyota
A state-space model is a statistical framework for inferring latent states from observed time-series data. However, inference with nonlinear and high-dimensional state-space models remains challenging. To this end, an approach based on diffusion models-a powerful class of deep generative models-has been developed, known as Score-based Data Assimilation (SDA)
Oscar Peralta
Telek (2022) asked whether a rational arrival process (RAP), specified by matrices ${G}_0$ and ${G}_1$ and an initial row vector ${\nu}$, with strictly positive joint densities and a unique dominant real eigenvalue of ${G}_0$ must admit an equivalent Markovian arrival process (MAP). A counterexample of order $3$ is given, showing the answer is no, and that t
Dogfight Search: A Swarm-Based Optimization Algorithm for Complex Engineering Optimization and Mountainous Terrain Path Planning
cs.AIYujing Sun, Jie Cai, Xingguo Xu, Yuansheng Gao
Dogfight is a tactical behavior of cooperation between fighters. Inspired by this, this paper proposes a novel metaphor-free metaheuristic algorithm called Dogfight Search (DoS). Unlike traditional algorithms, DoS draws algorithmic framework from the inspiration, but its search mechanism is constructed based on the displacement integration equations in kinem
Jae-Young Kang, Hoonhee Cho, Taeyeop Lee, Minjun Kang
Event cameras provide microsecond latency, making them suitable for 6D object pose tracking in fast, dynamic scenes where conventional RGB and depth pipelines suffer from motion blur and large pixel displacements. We introduce EventTrack6D, an event-depth tracking framework that generalizes to novel objects without object-specific training by reconstructing
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
cs.CVNanxi Li, Xiang Wang, Yuanjie Chen, Haode Zhang
While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-level physics reasoning. In this work, we investigate the firs
A Wide and Deep Exploration of Radio-detected Active Galactic Nuclei with Subaru HSC (WERGS). XIII. High-Redshift Radio Quasar candidates beyond Ultra-Steep Spectrum Selection: Dropout selection from HSC--VLASS over $\sim$1200 deg$^2$
astro-ph.GAYouwen Kong, Kohei Ichikawa, Hisakazu Uchiyama, Yuxing Zhong
We report the results of $g-$, $r-$, and $i-$dropout selections based on optical identifications of Very Large Array Sky Survey (VLASS) radio sources using the Hyper Suprime-Cam Subaru Strategic Program survey (HSC--SSP). By positional crossmatching within $1''.5$ between the VLASS Epoch~2 catalog and the HSC--SSP Wide-layer catalog ($i \lesssim 26$), we obt
Sangyi Wu, Junpu Guo, Xianghang Mi
Illicit online promotion is a persistent threat that evolves to evade detection. Existing moderation systems remain tethered to platform-specific supervision and static taxonomies, a reactive paradigm that struggles to generalize across domains or uncover novel threats. This paper presents a systematic study of In-Context Learning (ICL) as a unified framewor
Residue Constraints in the Rank-Three Lifting Problem for Projective-Plane Incidence Matrices
math.RAJaehwan Kim
We study the rank-three lifting problem for incidence matrices of finite projective planes through residue-level determinant constraints invisible to tropical valuations alone. In residue characteristic $\neq 3$, any rank-$\le 3$ lift of the incidence matrix of a projective plane of order $q \ge 3$ forces $\Omega(q^8)$ distinct admissible $2 \times 2$ zero r
M. Zeleny-Mora, R. Gaitán-Lozano, R. Martinez
We compute the complete set of one-loop contributions to the muon anomalous magnetic moment, $a_{\mu}=(g-2)_{\mu}/2$, in the Doublet Left-Right Symmetric Model (DLRSM), based on the gauge group $SU(2)_{L}\otimes SU(2)_{R}\otimes U(1)_{B-L}$ with neutrino masses generated via the inverse seesaw (ISS) mechanism. We evaluate all four one-loop topologies VFF, SF
Yakov Pyotr Shkolnikov
Deep learning training is non-deterministic: identical code with different random seeds produces models that agree on aggregate metrics but disagree on individual predictions, with per-class AUC swings exceeding 20 percentage points on rare clinical classes. We present a framework for verified bit-identical training that eliminates three sources of randomnes
Leonid Polterovich
In the contact-geometric approach to general relativity, the sky of an event - namely, the set of all incoming light rays - forms a Legendrian submanifold of the spherical cotangent bundle of a Cauchy hypersurface. When the hypersurface is chosen to be the Minkowski hyperboloid, a hyperbolic version of the hodograph transform identifies this bundle with a th
Gibbs measure for the HC-Blume-Capel model in the case of a "wand" type graph on a Cayley tree
math-phNosirjon M. Khatamov, Malika A. Kodirova
In this paper, we investigate translation-invariant splitting Gibbs measures (TISGMs) for the HC-Blume-Capel model on a "wand" graph embedded in the Cayley tree of arbitrary order $k \geq 2$. It is known that there is the exact critical value $\theta_{cr}$ such that for $\theta \geq \theta_{cr}$, there exists a unique TISGM, whereas for $\theta < \theta_{cr}
Rohan Pandey, Eric Ye, Michael Li
As Large Language Models (LLMs) achieve increasingly sophisticated performance on complex reasoning tasks, current architectures serve as critical proxies for the internal heuristics of frontier models. Characterizing emergent reasoning is vital for long-term interpretability and safety. Furthermore, understanding how prompting modulates these processes is e
Julio Candanedo, Alejandro Patiño
Diffusion maps (DMAP) are often used as a dimensionality-reduction tool, but more precisely they provide a spectral representation of the intrinsic geometry rather than a complete charting method. To illustrate this distinction, we study a Swiss roll with known isometric coordinates and compare DMAP, Isomap, and UMAP across latent dimensions. For each repres
Sangcheol Sim
Downlink bottlenecks motivate onboard systems that prioritize hazards without transmitting raw pixels. We study a strict setting where a ground station uplinks only compact embeddings plus metadata, and an onboard system performs vector search to triage new captures. We ask whether this embedding-only pipeline remains useful under explicit remote-sensing shi
Kexu Wang, Liangda Fang
Uniform interpolation is the property that, for any formula and set of atoms, there exists the strongest consequence omitting those atoms. It plays a central role in knowledge representation and reasoning tasks such as knowledge update and information hiding. This paper studies the uniform interpolation property in epistemic modal logics with distributed kno
RBF-Generated Finite Difference Method Coupled with Quadratic Programming for Solving PDEs on Surfaces with Derivative Boundary Conditions
math.NAPeng Chen, Shixiao Willing Jiang, Rongji Li, Qile Yan
Derivative boundary conditions introduce challenges for mesh-free discretizations of PDEs on surfaces, especially when the domain is represented by randomly sampled point clouds. The recently developed two-step tangent-space RBF-generated finite difference (RBF-FD) method provides high accuracy on closed surfaces. However, it may lose stability when applied
Zhuo Zhang, Lian-Bao Jia, Reyes J. F. Eduardo
In this work, we investigate a neutrinophilic low-mass dark matter model mediated by a pseudoscalar particle. Since dark matter lacks Standard Model gauge charges, new interactions are required to connect it to the visible sector. Traditional indirect detection searches for annihilation products, such as cosmic rays, become ineffective when the annihilation
Transfer Learning for an Endangered Slavic Variety: Dependency Parsing in Pomak Across Contact-Shaped Dialects
cs.CLSercan Karakaş
This paper presents new resources and baselines for Dependency Parsing in Pomak, an endangered Eastern South Slavic language with substantial dialectal variation and no widely adopted standard. We focus on the variety spoken in Turkey (Uzunk\"opr\"u) and ask how well a dependency parser trained on the existing Pomak Universal Dependencies treebank, which was
CARLA-Air: Fly Drones Inside a CARLA World -- A Unified Infrastructure for Air-Ground Embodied Intelligence
cs.ROTianle Zeng, Yanci Wen, Hong Zhang
The convergence of low-altitude economies, embodied intelligence, and air-ground cooperative systems creates growing demand for simulation infrastructure capable of jointly modeling aerial and ground agents within a single physically coherent environment. Existing open-source platforms remain domain-segregated: driving simulators lack aerial dynamics, while
Joseph M. Hellerstein
Classical complexity theory measures the cost of computing a function, but many computational tasks require committing to one valid output among several. We introduce determination depth -- the minimum number of sequential layers of irrevocable commitments needed to select a single valid output -- and show that no amount of computation can eliminate this cos
Arundhathi Dev, Justin Zhan
Optical character recognition remains critical infrastructure for document digitization, yet state-of-the-art performance is often restricted to well-resourced institutions by prohibitive computational barriers. End-to-end transformer architectures achieve strong accuracy but demand hundreds of GPU hours for domain adaptation, limiting accessibility for prac
Adapting SAM to Nuclei Instance Segmentation and Classification via Cooperative Fine-Grained Refinement
cs.CVJingze Su, Tianle Zhu, Jiaxin Cai, Zhiyi Wang
Nuclei instance segmentation is critical in computational pathology for cancer diagnosis and prognosis. Recently, the Segment Anything Model has demonstrated exceptional performance in various segmentation tasks, leveraging its rich priors and powerful global context modeling capabilities derived from large-scale pre-training on natural images. However, dire
Taeyun Roh, Suhyeong Park, Dongyoung Lee, Eunyeong Jo
Multimodal multiple-choice question answering (MCQA) provides a standardized and objectively measurable setting for evaluating vision-language models (VLMs). However, because the MCQA format incorporates the candidate choices into the input context, it introduces several unintended biases. Previous work has primarily focused on structural biases, such as pre
Fractional Modeling of Thermoelastic Fracture Behavior in a Cracked PZT-4 Strip under Transient Thermal Loading
cond-mat.mtrl-sciDiksha, Soniya Chaudhary, Pawan Kumar Sharma
This paper investigates the thermoelastic fracture response of a transversely isotropic piezoelectric strip containing a vertical insulated crack under transient thermal shock loading and pre-existing stress fields. The analysis is conducted within the framework of generalized fractional heat conduction using the Ezzat model, which incorporates thermal relax
Lizhe Wan, Jiaqi Yang
In this paper, we study the nonlinear modulational instability of two-dimensional hydroelastic Stokes waves in infinite depth. We first justify a focusing cubic nonlinear Schr\"odinger (NLS) approximation result for 2D deep hydroelastic wave system in the spirit of Ifrim-Tataru [22]. Then we exploit the instability mechanism of the cubic NLS to prove that th
Jiong Liu, Yingjie Xu, Xingcheng Zhou, Rui Song
Semantic segmentation across arbitrary sensor modalities faces significant challenges due to diverse sensor characteristics, and the traditional configurations for this task result in redundant development efforts. We address these challenges by introducing a universal arbitrary-modal semantic segmentation framework that unifies segmentation across multiple
Sai Rasmi Ranjan Mohanty, Priyank Vasu
We study a transformation surface associated with a zero mean curvature surface in the three-dimensional Heisenberg group with respect to two left-invariant semi-Riemannian metrics. We investigate the duality and prove that the transformation surface also has zero mean curvature. Furthermore, we derive the Sym formula for the dual surface in both metric case
Runkun Chen, Yixiong Fang, Pengyu Chang, Yuante Li
Deepfake speech detection systems are often limited to binary classification tasks and struggle to generate interpretable reasoning or provide context-rich explanations for their decisions. These models primarily extract latent embeddings for authenticity detection but fail to leverage structured acoustic evidence such as prosodic, spectral, and physiologica
Huimin Zeng, Yue Bai, Hailing Wang, Yun Fu
High dynamic range novel view synthesis (HDR-NVS) reconstructs scenes with dynamic details by fusing multi-exposure low dynamic range (LDR) views, yet it struggles to capture ambient illumination-dependent appearance. Implicitly supervising HDR content by constraining tone-mapped results fails in correcting abnormal HDR values, and results in limited gradien