Skip to content

April 2026 arXiv papers — page 160

Showing 15,90116,000 of 25,063 papers

  1. Hiroto Yamamoto, Katsuhiro Morita

    Accurately evaluating finite-temperature properties of quantum many-body systems remains a central challenge. Many existing quantum approaches typically require thermal-state preparation at each target temperature, making low-temperature calculations especially demanding in terms of circuit depth and accuracy. Here we introduce a distinct framework based onl

  2. Qian Zhang, Yuqin Cao, Yixuan Gao, Xiongkuo Min

    Video-to-Audio (V2A) generation is essential for immersive multimedia experiences, yet its evaluation remains underexplored. Existing benchmarks typically assess diverse audio types under a unified protocol, overlooking the fine-grained requirements of distinct audio categories. To address this gap, we propose VidAudio-Bench, a multi-task benchmark for V2A e

  3. Jia Li, Yu Zhang, Yin Chen, Zhenzhen Hu

    Facial action unit (AU) detection and facial expression (FE) recognition can be jointly viewed as affective facial behavior tasks, representing fine-grained muscular activations and coarse-grained holistic affective states, respectively. Despite their inherent semantic correlation, existing studies predominantly focus on knowledge transfer from AUs to FEs, w

  4. Chao Chen

    The BESIII collaboration has achieved important measurements in charmed purely leptonic and semi-leptonic decays using data samples collected at center-of-mass energies of 3.773 GeV, 4.128-4.226 GeV, and 4.237-4.669 GeV. This proceeding presents recent BESIII results on charmed purely leptonic and semileptonic decays, including measurements of branching frac

  5. Yuzhen Mao, Qitong Wang, Martin Ester, Ke Li

    Key-Value (KV) cache plays a crucial role in accelerating inference in large language models (LLMs) by storing intermediate attention states and avoiding redundant computation during autoregressive generation. However, its memory footprint scales linearly with sequence length, often leading to severe memory bottlenecks on resource-constrained hardware. Prior

  6. Viktoria Noel, Igor Lesanovsky

    We investigate the relaxation dynamics of a Rydberg gas in regimes where coherent processes and dissipation compete. In the strongly dissipative limit, the dynamics is known to be governed by an effective classical rate equation and to exhibit kinetically constrained, glassy relaxation towards a trivial stationary state. This behaviour originates from the Ry

  7. Gyaprasad, Rajneesh Joshi

    The propagation of the quantum states of light in dispersive and anisotropic media is a fundamental problem in quantum optics. We present a unified theoretical framework for the propagation of the quantum states of light in voltage-controlled nematic liquid crystals, incorporating both material dispersion and electrically tunable birefringence. By treating p

  8. Rongxiang Luo, Jiaqi Wen, Juncheng Guo

    Using nonequilibrium and equilibrium molecular dynamics simulations, we investigate heat conduction in a momentum-conserving mesoscopic fluid modeled by multiparticle collision dynamics. Across quasi-two-dimensional (q-2D) to three-dimensional (3D) systems, we identify three distinct transport regimes: (i) a \emph{ballistic regime}, where thermal conductivit

  9. Avi-ad Avraam Buskila

    Incorporating large language models (LLMs) in medical question answering demands more than high average accuracy: a model that returns substantively different answers each time it is queried is not a reliable medical tool. Online health communities such as Reddit have become a primary source of medical information for millions of users, yet they remain highl

  10. Tobias Mattsson, Samuel Nyberg, Anton Borg, Ricardo Britto

    The Model Context Protocol (MCP) is a new and emerging technology that extends the functionality of large language models, improving workflows but also exposing users to a new attack surface. Several studies have highlighted related security flaws, but MCP attack detection remains underexplored. To address this research gap, this study develops and evaluates

  11. Hung-Ting Su, Ting-Jun Wang, Jia-Fong Yeh, Min Sun

    Conventional Vision-and-Language Navigation (VLN) benchmarks assume instructions are feasible and the referenced target exists, leaving agents ill-equipped to handle false-premise goals. We introduce VLN-NF, a benchmark with false-premise instructions where the target is absent from the specified room and agents must navigate, gather evidence through in-room

  12. Jingkai Wang, Jue Gong, Zheng Chen, Kai Liu

    This paper provides a review of the NTIRE 2026 challenge on real-world face restoration, highlighting the proposed solutions and the resulting outcomes. The challenge focuses on generating natural and realistic outputs while maintaining identity consistency. Its goal is to advance state-of-the-art solutions for perceptual quality and realism, without imposin

  13. Jiahui Zhang, Rouyi Wang, Kuangqi Zhou, Tianshu Xiao

    Peptide therapeutics are widely regarded as the "third generation" of drugs, yet progress in peptide Machine Learning (ML) are hindered by the absence of standardized benchmarks. Here we present PepBenchmark, which unifies datasets, preprocessing, and evaluation protocols for peptide drug discovery. PepBenchmark comprises three components: (1) PepBenchData,

  14. Samuel Ferino, Rashina Hoda, John Grundy, Christoph Treude

    How software developers interact with Artificial Intelligence (AI)-powered tools, including Large Language Models (LLMs), plays a vital role in how these AI-powered tools impact them. While overreliance on AI may lead to long-term negative consequences (e.g., atrophy of critical thinking skills); underreliance might deprive software developers of potential g

  15. Hanming Fang, Xian Gu, Hanyin Yan, Wu Zhu

    We develop a high-precision classifier to measure artificial intelligence (AI) patents by fine-tuning PatentSBERTa on manually labeled data from the USPTO's AI Patent Dataset. Our classifier substantially improves the existing USPTO approach, achieving 97.0% precision, 91.3% recall, and a 94.0% F1 score, and it generalizes well to Chinese patents based on ci

  16. Zijia Lu, Jingru Yi, Jue Wang, Yuxiao Chen

    Referring multi-object tracking (RMOT) is a task of associating all the objects in a video that semantically match with given textual queries or referring expressions. Existing RMOT approaches decompose object grounding and tracking into separated modules and exhibit limited performance due to the scarcity of training videos, ambiguous annotations, and restr

  17. Tomoki Mihara

    Let $k$ be a complete valuation field. We formulate a free Banach $k$-vector space as a Banach $k$-vector space with an orthonormal Schauder basis, and an almost free Banach $k$-vector space as a non-Archimedean analogue of an almost free Abelian group. As non-Archimedean analogues of the classical facts that an almost free Abelian group is free under the as

  18. Xiaoyu Chen, Zhe Ju, Tianshun Miao, Yitong Yin

    We prove two results on the mixing times of Markov chains for two-spin systems. First, we show that the Glauber dynamics mixes in polynomial time for the Gibbs distributions of antiferromagnetic two-spin systems at the critical threshold of the uniqueness phase transition of the Gibbs measure on infinite regular trees. This completes the computational phase

  19. Yucheng Song, Chenxi Li, Haokang Ding, Zhining Liao

    In medical image segmentation across multiple modalities (e.g., MRI, CT, etc.) and heterogeneous data sources (e.g., different hospitals and devices), Domain Generalization (DG) remains a critical challenge in AI-driven healthcare. This challenge primarily arises from domain shifts, imaging variations, and patient diversity, which often lead to degraded mode

  20. BESIII Collaboration, M. Ablikim, M. N. Achasov, P. Adlarson

    Using $(2712.4\pm14.3)\times 10^6$ $\psi(3686)$ events collected with the BESIII detector operating at BEPCII, the branching fractions of $\chi_{cJ}\to\pi^+\pi^-\pi^0\pi^0$ ($J=0,~1,~2$) are measured via the radiative transition $\psi(3686)\to\gamma\chi_{cJ}$. The results are $\mathcal{B}(\chi_{c0} \to \pi^{+}\pi^{-}\pi^{0}\pi^{0}) = (3.10 \pm 0.01 \pm 0.14)

  21. Mengieong Hoi, Zhedong Zheng, Ping Liu, Wei Liu

    Deepfake content on social networks is increasingly produced through multiple \emph{sequential} edits to biometric data such as facial imagery. Consequently, the final appearance of an image often reflects a latent chain of operations rather than a single manipulation. Recovering these editing histories is essential for visual provenance analysis, misinforma

  22. Ivan Kitov

    Seismic noise with an amplitude higher than that of the sought signal is a challenge for detection. Several techniques have been developed to suppress the ambient noise and to reduce the detection threshold in order to find signals with the lowest possible amplitudes produced by events with the magnitudes significant for scientific research and technical app

  23. Suyoung Bae, CheolWon Na, Jaehoon Lee, Yumin Lee

    As Large Language Models (LLMs) have become capable of generating long and descriptive code summaries, accurate and reliable evaluation of factual consistency has become a critical challenge. However, previous evaluation methods are primarily designed for short summaries of isolated code snippets. Consequently, they struggle to provide fine-grained evaluatio

  24. Iris Groher, Patrick Heissenberger, Michael Vierhauser

    Large Language Models (LLMs) have become part of how students solve programming tasks, offering immediate explanations and even full solutions. Previous work has highlighted that novice programmers often heavily rely on LLMs, thereby neglecting their own problem-solving skills. To address this challenge, we designed a course-specific online Python tutor that

  25. Xiaoda Yang, Yuxiang Liu, Shenzhou Gao, Can Wang

    Modern vision-language models achieve strong performance in static perception, but remain limited in the complex spatiotemporal reasoning required for embodied, egocentric tasks. A major source of failure is their reliance on temporal priors learned from passive video data, which often leads to spatiotemporal hallucinations and poor generalization in dynamic

  26. Xinyi Huang

    Selecting the right knowledge is critical when using large language models (LLMs) to solve domain-specific data analysis tasks. However, most retrieval-augmented approaches rely primarily on lexical or embedding similarity, which is often a weak proxy for the task-critical knowledge needed for multi-step reasoning. In many such tasks, the relevant knowledge

  27. Yifan He, Xuchen Hua, Lei Shen, Kai Zhang

    The topological confinement is a new mechanism that allows the transmission of cutoff orbital angular momentum (OAM) modes with negligible loss in ring-core fibers (RCFs) and provides a natural immunity against mode coupling. We investigate the influence of fiber design parameters and wavelength on the characteristics of topologically confined modes (TCMs) i

  28. Lincoln Spencer, Song Wang, Chen Chen

    Surgical phase segmentation is central to computer-assisted surgery, yet robust models remain difficult to develop when labeled surgical videos are scarce. We study data-efficient phase segmentation for manual small-incision cataract surgery (SICS) through a controlled comparison of visual representations. To isolate representation quality, we pair each visu

  29. Roi Ben-Gigi, Yuval David, Fabiana Fournier, Lior Limonad

    AI agent development relies heavily on natural language prompting to define agents' tasks, knowledge, and goals. These prompts are interpreted by Large Language Models (LLMs), which govern agent behavior. Consequently, agentic performance is susceptible to variability arising from imprecise or ambiguous prompt formulations. Identifying and correcting such is

  30. Chenhan Jiang, Yu Chen, Qingwen Zhang, Jifei Song

    The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse and precise camera trajectories. While real-world captures are photorealistic, they are typically sparse and discrete. Conversely, synthetic data scales but suffers from a domain gap and often lacks realistic

  31. Hu Ligui, Meng Qingxin, Tang Maoning

    This paper focuses on the discrete-time backward stochastic linear quadratic (BSLQ) optimal control problem with nonhomogeneous system terms and cost function cross terms. The terminal constraint of such systems distinguishes it from forward stochastic systems, posing unique challenges for analysis and solution. Within the Hilbert space framework, we first c

  32. Linjie Zhao

    We study the symmetric simple exclusion process with Glauber dynamics. When the process starts from a nonequilibrium measure, we prove central limit theorems for the occupation time in dimension two, and sample path moderate deviation principles in dimension one. For the fluctuations, we use the martingale method and the sharp relative entropy method from [J

  33. Johin Johny Arimbur

    Large language models frequently fail to produce correct code on their first attempt, yet most benchmarks evaluate them in a single-shot setting. We investigate iterative self-repair (feeding execution errors back to the model for correction) across seven models spanning three families and both open-weight and proprietary providers: Llama 3.1 8B, Llama 3.3 7

  34. Danni Liu, Bo Liu, Yuxin Hu, Hantao Zhao

    Psychological client simulators have emerged as a scalable solution for training and evaluating counselor trainees and psychological LLMs. Yet existing simulators exhibit unrealistic over-compliance, leaving counselors underprepared for the challenging behaviors common in real-world practice. To bridge this gap, we present ResistClient, which systematically

  35. Xiaoda Yang, Shuai Yang, Can Wang, Jingyang Xue

    Vision-Language Models (VLMs) have made significant strides in static image understanding but continue to face critical hurdles in spatiotemporal reasoning. A major bottleneck is "multi-image reasoning hallucination", where a large performance drop between forward and reverse temporal queries reveals a dependence on superficial shortcuts instead of s

  36. M. Burgess

    Agent based systems are more common than we may think. A Promise Theory perspective on cooperation, in systems of human-machine agents, offers a unified perspective on organization and functional design with semi-automated efforts, in terms of the abstract properties of autonomous agents, This applies to human efforts, hardware systems, software, and artific

  37. Bingzhe Wu, Haotian Lu, Yuchen Mou

    Current large language models (LLMs), even those explicitly trained for reasoning, often struggle with ambiguous content moderation cases due to misleading "decision shortcuts" embedded in context. Inspired by cognitive psychology insights into expert moderation, we introduce \caro (Chain-of-Analogy Reasoning Optimization), a novel two-stage training framewo

  38. Shivam Chauhan, Ajay Pundhir

    Modern audio systems universally employ mel-scale representations derived from 1940s Western psychoacoustic studies, potentially encoding cultural biases that create systematic performance disparities. We present a comprehensive evaluation of cross-cultural bias in audio front-ends, comparing mel-scale features with learnable alternatives (LEAF, SincNet) and

  39. Haotian Lu, Yuchen Mou, Bingzhe Wu

    Content moderation in online platforms faces persistent challenges due to the evolving complexity of user-generated content and the limitations of traditional rule-based and machine learning approaches. While recent advances in large language models (LLMs) have enabled more sophisticated moderation via direct prompting or fine-tuning, these approaches often

  40. Saket Jha, Karthikeya S. M. Yelisetty, Singabattu Sathya, Shamik Sural

    Recent advances in research on Attribute-based Access Control (ABAC) has led to the development of several ingenious methods for representing and enforcing organizational security policies. However, so far little effort has been spent towards building a tool for generating large-scale synthetic datasets that can be used to test the developed ABAC systems. In

  41. Hongxi Mao, Wei Zhou, Mengting Jia, Tao Fang

    Machine learning for tabular data remains constrained by poor schema generalization, a challenge rooted in the lack of semantic understanding of structured variables. This challenge is particularly acute in domains like clinical medicine, where electronic health record (EHR) schemas vary significantly. To solve this problem, we propose Schema-Adaptive Tabula

  42. Yudong Han, Yong Wang, Zaiquan Yang, Zhen Qu

    Multimodal latent reasoning has emerged as a promising paradigm that replaces explicit Chain-of-Thought (CoT) decoding with implicit feature propagation, simultaneously enhancing representation informativeness and reducing inference latency. By analyzing token-level gradient dynamics during latent training, we reveal two critical observations: (1) visual tok

  43. Tao Liu, Xuehe Wang

    Federated learning has become a popular paradigm for privacy protection and edge-based machine learning. However, defending against differential attacks and devising incentive strategies remain significant bottlenecks in this field. Despite recent works on privacy-aware incentive mechanism design for federated learning, few of them consider both data volume

  44. Sheng Chen, Kai Huang

    Given a finite extension $K/k$ of number fields and a smooth quasi-projective variety $X$ over $K$. If the abelianized fundamental group of $X$ is trivial, we prove that there is a natural identification between Brauer-Manin sets of $X$ and its Weil restriction $R_{K/k}X$. If $X$ is projective and $Pic(X\times_{K}\overline{k})$ is a torsion-free abelian grou

  45. Xiangyang Yin, Xingyu Liu, Tianhua Xia, Bo Bao

    Outliers have emerged as a fundamental bottleneck in preserving accuracy for low-precision large models, particularly within Mixture-of-Experts (MoE) architectures that are increasingly central to large-scale language modeling. Under post-training quantization (PTQ), these outliers induce substantial quantization errors, leading to severe accuracy degradatio

  46. Jie Zhang, Jiapeng Guan, Hao Zhou, Xiaomeng Han

    Block Floating-Point (BFP) is emerging as an attractive data format for edge Neural Processing Units (NPUs), combining wide dynamic range with high hardware efficiency. However, its behavior under hardware faults and suitability for safety-critical deployments remain underexplored. Here, we present the first in-depth empirical reliability study of BFP-based

  47. Mahir Labib Dihan, Md Ashrafur Rahman Khan

    Automating real-world software engineering tasks remains challenging for large language model (LLM)-based agents due to the need for long-horizon reasoning over large, evolving codebases and making consistent decisions across interdependent actions. Existing approaches typically rely on static prompting strategies or handcrafted heuristics to select actions

  48. Hyunyoung Han, Murad Eynizada, Son Xuan Nghiem, Sang Ho Yoon

    Online dance tutorials have gained widespread popularity. However, many novices encounter difficulties when dance motion complexity exceeds their skill level, potentially leading to discouragement. This study explores dance motion simplification to address this challenge. We surveyed 30 novices to identify challenging movements, then conducted focus groups w

  49. Oren Goldberg, Noa Mazurski, Uriel Levy

    Mie-void metasurfaces have so far been developed mainly in reflection, where subwavelength voids embedded in high-index media support localized resonances and spectrally selective optical responses. Yet, many optical systems could benefit from integrating such optical elements operating in transmission mode. Motivated by this great need, we hereby introduce

  50. Linn Evenseth, Kamil Galewski, Witold Jarnicki, Piero Lafiosca

    We present a computational platform for modeling chemical reactions in complex molecular environments, focused on ligand-protein binding in drug discovery. The platform implements our new quantum-in-quantum-in-classical (QM/QM/MM) multiscale embedding model that integrates molecular dynamics with a quantum-information-enhanced density matrix embedding theory

  51. Andrea Steinfurth, Sebastian Weidemann, Julia Görsch, Tom Sheppard

    Time is the odd dimension out: Unlike space, it follows the arrow of time, forbidding back-reflections and requiring momentum yet not energy conservation. Tailored temporal variations manipulate momentum bands and engineer waves in time. We show that momentum bands exhibit unique topology, hidden when conventionally considering energy bands: Complex momentum

  52. Haopeng Chen, Yihao Ai, Kabeen Kim, Robby T. Tan

    Low-visibility scenarios, such as low-light conditions, pose significant challenges to human pose estimation due to the scarcity of annotated low-light datasets and the loss of visual information under poor illumination. Recent domain adaptation techniques attempt to utilize well-lit labels by augmenting well-lit images to mimic low-light conditions. But han

  53. Jiapeng Guan, Jie Zhang, Hao Zhou, Ran Wei

    DNNs and LLMs increasingly rely on hardware accelerators, including in safety-critical domains, while technology scaling and growing model complexity make hardware faults more frequent. Existing system-level mechanisms typically treat the NPU as a monolithic unit, using coarse-grained replication that incurs prohibitive performance and hardware overheads, le

  54. Doried Ghader, Bilal Jabakhanji

    We investigate magnon-phonon (MP) excitations in a Neel-type two-dimensional ferromagnetic skyrmion crystal (SkX) stabilized on a triangular spin lattice by Dzyaloshinskii-Moriya interaction (DMI). Although the lowest two magnon bands of the bare SkX are topologically trivial, we show that coupling to lattice vibrations reconstructs the low-energy sector and

  55. Shuaida He, Yangzhou Chen, Xin Chen

    Modern regression analysis often involves responses and predictors taking values in the same or distinct metric spaces. To rank non-Euclidean heterogeneous predictors in regression by explanatory strength, analogous to the classical $R^2$, we introduce the Fr\'echet correlation coefficient (FCC), defined as the relative reduction in the Fr\'echet variance of

  56. Mahir Labib Dihan, Faria Binta Awal, Md. Ishrak Ahsan

    Retrieving the correct set of files from a large codebase is a crucial step in Automated Program Repair (APR). High recall is necessary to ensure that the relevant files are included, but simply increasing the number of retrieved files introduces noise and degrades efficiency. To address this tradeoff, we propose PatchRecall, a hybrid retrieval approach that

  57. Yu Li, Xiaoran Shang, Qizhi Pei, Yun Zhu

    Post-training data plays a pivotal role in shaping the capabilities of Large Language Models (LLMs), yet datasets are often treated as isolated artifacts, overlooking the systemic connections that underlie their evolution. To disentangle these complex relationships, we introduce the concept of \textbf{data lineage} to the LLM ecosystem and propose an automat

  58. Isaac M Hair, Amit Sahai

    We give a public key encryption scheme with plausible quasi-exponential security based on the conjectured intractability of two constraint satisfaction problems (CSPs), both of which are instantiated with a corruption rate of $1 - o(1)$. First, we conjecture the hardness of a new large alphabet random predicate CSP (LARP-CSP) defined over an arbitrary but st

  59. Hongxing Chen, Xiaohu Chen, Jinbi Zhang

    In a general triangulated category, the finiteness of the finitistic dimension serves as a prerequisite for a categorical obstruction, via the singularity category, to the existence of bounded $t$-structures. In this paper, we investigate the finitistic, big finitistic, and global dimensions, and establish explicit inequalities that relate these dimensions o

  60. Riccardo Marchesi

    The ubiquitous $3/4$ metabolic scaling exponent, known as Kleiber's law, has long been attributed to the minimization of viscous dissipation within fractal transport networks. In this paper, we invert this standard narrative, demonstrating that Kleiber's law is fundamentally a signature of pulsatile wave physics rather than steady-state geometry. By coupling

  61. Laurent Delisle, Amine Jaouadi

    We investigate the integrable structure and soliton dynamics of a coupled modified Korteweg-de Vries (cmKdV) system with a real symmetric coupling matrix. We introduce a vector reformulation of Hirota's bilinear formalism in which both the bilinear equations and their solutions are expressed directly at the vector level, rather than through a component-wise

  62. Di Kevin Gao, Jingdao Chen, Shahram Rahimi

    As artificial intelligence (AI) systems grow more powerful, autonomous, and embedded in critical infrastructure, their identification and traceability become foundational to regulatory oversight and sustainable digital governance. In digitally transformed enterprises, long-term sustainability depends on transparent, accountable, and lifecycle-governed AI sys

  63. Shinichiro Kakuta

    We study the volume conjecture of the colored Jones invariants with sequences of colors corresponding to the deformation of the hyperbolic structure of a link complement. In particular, we investigate certain limits of the colored Jones invariants of the figure-eight knot and the Borromean rings and show that the limits are related to the volumes of hyperbol

  64. Molla Basir Ahamed, Rajesh Hossain

    This paper investigates the geometric and analytical properties of harmonic mappings $f$ in the unit disk $\mathbb{D}$ induced by boundary functions $F$ belonging to the Lebesgue spaces $L^{p}(\mathbb{T})$ for $1 \le p \le \infty$. We first establish a sharp Bohr-type inequality for the class of bounded harmonic mappings. Specifically, we prove that for a fi

  65. Guowen Li, Yuepeng Zhang, Shunyu Zhang, Yi Zhang

    Large-scale short-video search ranking models are typically trained on sparse co-occurrence signals over hashed item identifiers (HIDs). While effective at memorizing frequent interactions, such ID-based models struggle to generalize to long-tailed items with limited exposure. This memorization-generalization trade-off remains a longstanding challenge in suc

  66. Mingfei Lu, Yi Zhang, Mengjia Wu, Yue Feng

    Legal consultation question answering (Legal CQA) presents unique challenges compared to traditional legal QA tasks, including the scarcity of high-quality training data, complex task composition, and strong contextual dependencies. To address these, we construct JurisCQAD, a large-scale dataset of over 43,000 real-world Chinese legal queries annotated with

  67. Ye Su, Mingrui Ye, Yining Wang, Jipeng Guo

    Standard resampling ratios (e.g., $\alpha \approx 0.632$) are widely used as default baselines in ensemble learning for three decades. However, how these ratios interact with a base learner's intrinsic functional complexity in finite samples lacks a exact mathematical characterization. We leverage the Hoeffding-ANOVA decomposition to derive the first exact,

  68. Dipanwita Ghoshal

    Understanding dynamical facilitation in nonequilibrium glass-forming systems driven by active forces remains an open challenge. In particular, it is unclear whether facilitation survives in active glasses, where persistent self-propulsion breaks detailed balance and introduces directional memory. Here, we use large-scale simulations of a two-dimensional athe

  69. Rajat Adak

    A hypergraph $H$ is said to be \emph{linear} if every pair of vertices lies in at most one hyperedge. Given a family $\mathcal{F}$ of $r$-uniform hypergraphs (also called $r$-graphs), an $r$-graph $H$ is said to be \emph{$\mathcal{F}$-free} if it contains no member of $\mathcal{F}$ as a subhypergraph. The \emph{linear Tur\'{a}n number} $ex_r^{\mathrm{lin}}(n

  70. Arjun Somayazulu, Kristen Grauman

    Visual feedback is critical for motor skill acquisition in sports and rehabilitation, and psychological studies show that observing near-perfect versions of one's own performance accelerates learning more effectively than watching expert demonstrations alone. We propose to enable such personalized feedback by automatically editing a person's motion t

  71. Ya-Feng Lo, Dmitrii Kobylianskii, Benjamin Nachman, Eilam Gross

    We present the application of Parnassus, a generative model for full detector simulation and reconstruction, to the ALEPH detector at the Large Electron-Positron Collider (LEP). Training on simulated $e^+e^-$ to Z to qqbar events processed through the ALEPH detector simulation and reconstruction, we demonstrate that Parnassus faithfully reproduces the detect

  72. Candi Zheng, Yuan Lan

    Diffusion models are often introduced from multiple perspectives, such as VAEs, score matching, or flow matching, accompanied by dense and technically demanding mathematics that can be difficult for beginners to grasp. One classic question is: how does the reverse process invert the forward process to generate data from pure noise? This article systematicall

  73. Yuerang Li, Zipeng Wang

    We establish a converse of the Shimorin--Pel\'{a}ez--R\"{a}tty\"{a}--Wick theorem. Specifically, we obtain necessary and sufficient conditions for a Shimorin kernel to be the kernel of a radial, logarithmically subharmonic weighted Bergman space.

  74. Sonali Gupta, Kiran Bajar, Alexander McFarland, Amit Kumar

    Optical imaging of quantum emitters is essential for a wide range of quantum applications. Conventional confocal imaging relies on point-by-point raster scanning, which is inherently time-consuming and photon-inefficient, particularly for sparse emitter distributions and photon-limited samples. Here, we demonstrate a compressive sensing-based imaging approac

  75. Arvid Siqveland

    We prove that the classical algebraic varieties over algebraically closed fields can be defined over arbitrary fields $k.$ Then we prove that for associative algebras $A$, there exist local representing objects $A_M$ for simple modules $M.$ Replacing the localization in maximal ideals in the commutative situation with the local representations in simple modu

  76. Qiyang Chen, Guozheng Li, Xingqi Wang, Gerile Aodeng

    Hierarchical tables are an important structure for organizing data with inherent hierarchical relationships. Existing studies have extensively explored methods for data fact exploration from tabular data. In particular, some studies have directly integrated visual data facts into the original table structure to support in-situ exploration, because embedding

  77. Xinlei Guan, David Arosemena, Tejaswi Dhandu, Kuan Huang

    The rapid growth of generative AI has introduced new challenges in content moderation and digital forensics. In particular, benign AI-generated images can be paired with harmful or misleading text, creating difficult-to-detect misuse. This contextual misuse undermines the traditional moderation framework and complicates attribution, as synthetic images typic

  78. Qingyang Li

    The exponential growth of user-generated movie reviews on digital platforms has made accurate text sentiment classification a cornerstone task in natural language processing. Traditional models, including standard BERT and recurrent architectures, frequently struggle to capture long-distance semantic dependencies and resolve ambiguous emotional expressions i

  79. Naichuan Zheng, Hailun Xia, Zepeng Sun, Weiyi Li

    Wearable IMU-based Human Activity Recognition (HAR) relies heavily on Deep Neural Networks (DNNs), which are burdened by immense computational and buffering demands. Their power-hungry floating-point operations and rigid requirement to process complete temporal windows severely cripple battery-constrained edge devices. While Spiking Neural Networks (SNNs) of

  80. Jiarui Guan, Wenshuai Zhao, Zhengtao Zou, Juho Kannala

    Video variational autoencoders (VAEs) used in latent diffusion models typically require a sufficiently large number of latent channels to ensure high-quality video reconstruction. However, recent studies have revealed that an excessive number of latent channels can impede the convergence of latent diffusion models and deteriorate their generative performance

  81. Songtao Mao

    Noisy $k$-XOR is a basic average-case inference problem in which one observes random noisy $k$-ary parity constraints and seeks to recover, or more weakly, detect, a hidden Boolean assignment. A central question is to characterize the tradeoff among sample complexity, noise level, and running time. We give a recovery algorithm, and hence also a detection alg

  82. Peixuan Zhang, Chang Zhou, Ziyuan Zhang, Hualuo Liu

    The surging demand for adapting long-form cinematic content into short videos has motivated the need for versatile automatic video compilation systems. However, existing compilation methods are limited to predefined tasks, and the community lacks a comprehensive benchmark to evaluate the cinematic compilation. To address this, we introduce CineBench, the fir

  83. Hengyu Zhang, Xuyun Zhang, Pengxiang Zhan, Linhao Luo

    Recent advances in large language models (LLMs) have enabled promising progress in diagnosis prediction from electronic health records (EHRs). However, existing LLM-based approaches tend to overfit to historically observed diagnoses, often overlooking novel yet clinically important conditions that are critical for early intervention. To address this, we prop

  84. Shi Chen, Xuecheng Wu, Heli Sun, Yunyun Shi

    Affective Image Manipulation (AIM) aims to evoke specific emotions through targeted editing. Current image editing benchmarks primarily focus on object-level modifications in general scenarios, lacking the fine-grained granularity to capture affective dimensions. To bridge this gap, we introduce the first benchmark designed for AIM termed AIM-Bench. This ben

  85. Noha Hassan, Xavier Fernando, Halim Yanikomeroglu

    As a key enabler for sixth-generation (6G) wireless communications, reconfigurable intelligent surfaces (RISs) provide the flexibility to control signal strength. Nevertheless, optimizing hundreds of elements is computationally expensive. To overcome this challenge, we present a quantum framework (QGCN) to jointly optimize the physical and electromagnetic re

  86. Sanjaya Poudel, Nikita Kunwor, Raj Simkhada, Mustafa Munir

    Despite recent advancements in the field of medical image analysis with the use of pretrained foundation models, the issue of distribution shifts between cross-source images largely remains adamant. To circumvent that issue, investigators generally train a separate model for each source. However, this method becomes expensive when we fully fine-tune pretrain

  87. Yige Yang, Man Zhang, Tao Yue

    Test optimization contains test case selection and minimization, which is an important challenge in software testing and has been addressed with search-based approaches intensively in the past. Inspired by the recent advancement of using quantum optimization solutions for addressing test optimization problems, we looked into Coherent Ising Machines (CIM), wh

  88. Qingyao Li, Weiwen Liu, Weinan Zhang, Yong Yu

    Recent advancements in Large Language Models (LLMs) have successfully employed search-based strategies to enhance code generation. However, existing methods typically rely on static, sparse public test cases for verification, leading to pseudo-correctness -- where solutions overfit the visible public tests but fail to generalize to hidden test cases. We argu

  89. Bo Li, Mingda Wang, Shikun Zhang, Wei Ye

    Instruction tuning relies on large instruction-response corpora whose quality and composition strongly affect downstream performance. We propose Answer Divergence-Guided Selection (ADG), which selects instruction data based on the geometric structure of multi-sample outputs. ADG draws several high-temperature generations per instruction, maps responses into

  90. Vibodha Bandara, Jordan D. Roberts, Dustin Keller

    The SpinQuest experiment at Fermilab employs a dynamically polarized solid ammonia target to probe the spin structure of the proton, requiring stable, optimized microwave-driven Dynamic Nuclear Polarization (DNP) under high radiation conditions. We present the design, operation, and automation of a 140 GHz microwave system based on an extended interaction os

  91. Dongbin Li, Alexander E. Litvak, Tingzhou Yu

    Let $M_n$ be an $n\times n$ random matrix with entries in $\{0, 1\}$, where each row is independently and uniformly sampled from the set of all vectors in $\{0, 1\}^n$ containing exactly $d$ ones. we establish quantitative lower bounds on the smallest singular value of the shifted matrices $M_n-z \mathbf{I}_n$ whenever $|z| \leq \sqrt{d}\, \log\log d$ and $

  92. Anamitra Ghorui, Aditi Raste, Uday P. Khedker

    Precise pointer analysis is a foundational component of many client analyses and optimizations. Scaling flow- and context-sensitive pointer analysis has been a long-standing challenge, suffering from combinatorial growth in both memory usage and runtime. Existing approaches address this primarily by reducing the amount of information tracked often, at the co

  93. BESIII Collaboration, M. Ablikim, M. N. Achasov, P. Adlarson

    We perform the first amplitude analysis of the singly Cabibbo-suppressed decays $D^+ \to \pi^+ \pi^{+(0)} \pi^{-(0)} \eta$, using $e^+e^-$ collision data collected with the BESIII detector at the center-of-mass energy of 3.773\,GeV, corresponding to an integrated luminosity of 20.3 $\rm{fb}^{-1}$. The absolute branching fractions of the $D^+ \to \pi^+ \pi^+

  94. Sara Fish, Yannai A. Gonczarowski, Jason Z. Tang, Salil Vadhan

    The differentially private (DP) facility location problem seeks to determine a socially optimal placement for a public facility while ensuring that each participating agent's location remains private. To privatize its input data, a DP mechanism must inject noise into its output distribution, producing a placement that will have lower expected social welfare

  95. Peixuan Zhang, Zijian Jia, Ziqi Cai, Shuchen Weng

    Effective poster design requires rapidly capturing attention and clearly conveying messages. Inspired by the ``contrast effects'' principle, we propose ReContraster, the first training-free model to leverage regional contrast to make posters stand out. By emulating the cognitive behaviors of a poster designer, ReContraster introduces the compositional multi-

  96. Sina Mansouri, Mohit Marvania, Vibhavari Ashok Shihorkar, Han Ngoc Tran

    Medical large language models are typically evaluated on idealized patient cases that do not reflect how real patients communicate. We introduce VeriSim, a patient simulation framework that injects controllable noise along six clinically grounded communication dimensions while substantially preserving each patient's medical record. Truth adherence is sup

  97. Zhenwei Li, Dandan Wei, Shi Jia, Hailiang Chen

    The massive binary common envelope (CE) phase plays a pivotal role in the formation of close black hole/neutron star (BH/NS) binaries, yet significant uncertainties remain in our understanding of this process. In this study, we aim to constrain the massive binary CE phase by systematically reconstructing three observed BH X-ray binaries (BHXBs): GRO J1655-40

  98. Ziheng Guo, Danqun Zheng, Shuai Li, Chengwei Chen

    Purpose: Deep learning-based MRI artifact correction methods often demonstrate poor generalization to clinical data. This limitation largely stems from the inability of deep learning models in reliably distinguishing motion artifacts from true anatomical structures, due to insufficient awareness of artifact characteristics. To address this challenge, we prop

  99. Jielin Qiu, Ming Zhu, Wenting Zhao, Zhiwei Liu

    Audio-native large language models (audio-LLMs) commonly use Whisper as their audio encoder. However, Whisper was trained exclusively on speech data, producing weak representations for music and environmental sound. This forces downstream audio-LLMs to compensate through extensive training on large-scale non-speech data. We present Whisper-AuT, a domain-adap

  100. Chenyu Wang, Weicheng Dai, Han Liu, Wenchao Li

    Vision--language models (VLMs) for radiology report generation (RRG) can produce long-form chest CT reports from volumetric scans and show strong potential to improve radiology workflow efficiency and consistency. However, existing methods face two key limitations: (i) training supervision is often coarse, aligning a whole CT volume with a full free-text rep