May 2025 arXiv papers — page 53
Showing 5,201–5,300 of 24,552 papers
Monocle: Hybrid Local-Global In-Context Evaluation for Long-Text Generation with Uncertainty-Based Active Learning
cs.CLXiaorong Wang, Ting Yang, Zhu Zhang, Shuo Wang
Assessing the quality of long-form, model-generated text is challenging, even with advanced LLM-as-a-Judge methods, due to performance degradation as input length increases. To address this issue, we propose a divide-and-conquer approach, which breaks down the comprehensive evaluation task into a series of localized scoring tasks, followed by a final global
G. E. Volovik
Recently the difference between the Gibbons-Hawking temperature $T_{\rm GH}$ attributed to the Hawking radiation from the de Sitter cosmological horizon and the twice as high local temperature of the de Sitter state, $T=H/\pi=2T_{\rm GH}$, has been discussed by Hughes and Kusmartsev from the topological point of view (see arXiv:2505.05814). According to thei
Thomas Montandon, Elsa M. Teixeira, Adèle Poudou, Vivian Poulin
Decaying dark matter (DDM) has emerged as an interesting framework to extend the $\Lambda$-cold-dark-matter (LCDM) model, as many particle physics models predict that dark matter may not be stable over cosmic time and can impact structure formation. In particular, a model in which DDM decays at a rate $\Gamma$ and imprints a velocity kick $v$ onto its decay
FunReason: Enhancing Large Language Models' Function Calling via Self-Refinement Multiscale Loss and Automated Data Refinement
cs.LGBingguang Hao, ZengZhuang Xu, Maolin Wang, Yuntao Wen
The integration of large language models (LLMs) with function calling has emerged as a crucial capability for enhancing their practical utility in real-world applications. However, effectively combining reasoning processes with accurate function execution remains a significant challenge. Traditional training approaches often struggle to balance the detailed
Claudio Dappiaggi, Vincenzo Morinelli, Gerardo Morsella, Alessio Ranallo
We consider a four-dimensional globally hyperbolic spacetime $(M,g)$ conformal to Minkowski spacetime, together with a massless, conformally coupled scalar field. Using a bulk-to-boundary correspondence, one can establish the existence of an injective $*$-homomorphism $\Upsilon_M$ between $\mathcal{W}(M)$, the Weyl algebra of observables on $M$ and a counter
Tonmoy Hasan, Razvan Bunescu
The affective attitude of liking a recommended item reflects just one category in a wide spectrum of affective phenomena that also includes emotions such as entranced or intrigued, moods such as cheerful or buoyant, as well as more fine-grained affective states, such as "pleasantly surprised by the conclusion". In this paper, we introduce a novel recommendat
Syamantak Kumar, Daogao Liu, Kevin Tian, Chutong Yang
Estimating the geometric median of a dataset is a robust counterpart to mean estimation, and is a fundamental problem in computational geometry. Recently, [HSU24] gave an $(\varepsilon, \delta)$-differentially private algorithm obtaining an $\alpha$-multiplicative approximation to the geometric median objective, $\frac 1 n \sum_{i \in [n]} \|\cdot - \mathbf{
Zhenzhen Song, Ziwei Liu, Hongji Li
Aiming at the problems of cross-modal feature fusion, low efficiency of long text modeling and lack of hierarchical semantic coherence in patent text semantic mining, this study proposes HGM-Net, a deep learning framework that integrates Hierarchical Comparative Learning (HCL), Multi-modal Graph Attention Network (M-GAT) and Multi-Granularity Sparse Attentio
J. Philip Haupt, Evelin M. C. Christlmaier, Pablo López Ríos, Nikolay A. Bogdanov
We apply the transcorrelated method to problems of multireference character. For this, we show that the choice of reference wavefunction during the Jastrow optimisation procedure is vital, and we propose a workflow wherein we use conventional multi-configurational methods to provide a reference wavefunction for Jastrow factor optimisation. This Jastrow funct
Eradicating the Unseen: Detecting, Exploiting, and Remediating a Path Traversal Vulnerability across GitHub
cs.CRJafar Akhoundali, Hamidreza Hamidi, Kristian Rietveld, Olga Gadyatskaya
Vulnerabilities in open-source software can cause cascading effects in the modern digital ecosystem. It is especially worrying if these vulnerabilities repeat across many projects, as once the adversaries find one of them, they can scale up the attack very easily. Unfortunately, since developers frequently reuse code from their own or external code resources
Yongan Yu, Mengqian Wu, Yiran Lin, Nikki G. Lobczowski
Assessing higher-order thinking skills in large language models (LLMs) remains a fundamental challenge, especially in tasks that go beyond surface-level accuracy. In this work, we propose THiNK (Testing Higher-order Notion of Knowledge), a multi-agent, feedback-driven evaluation framework grounded in Bloom's Taxonomy. THiNK frames reasoning assessment as an
Karolina Gorna, Nicolas Iooss, Yannick Seurin, Rida Khatoun
The widespread adoption of the Go programming language in infrastructure backends and blockchain projects has heightened the need for improved security measures. Established techniques such as unit testing, static analysis, and program fuzzing provide foundational protection mechanisms. Although symbolic execution tools have made significant contributions, o
Shubham Gandhi, Atharva Naik, Yiqing Xie, Carolyn Rose
We study cost-efficient collaboration between strong and weak language models for repository-level code generation, where the weak model handles simpler tasks at lower cost, and the most challenging tasks are delegated to the strong model. While many works propose architectures for this task, few analyze performance relative to cost. We evaluate a broad spec
Maurice Chiodo, Dennis Müller
The increasing deployment of Artificial Intelligence (AI) and other autonomous algorithmic systems presents the world with new systemic risks. While focus often lies on the function of individual algorithms, a critical and underestimated danger arises from their interactions, particularly when algorithmic systems operate without awareness of each other, or w
Sylvain Fichet, Rodrigo Fresneda, Lucas de Souza, Dmitri Vassilevich
Computing the vacuum expectation of fermion number operator on a soliton background is often challenging. A recent proposal in arXiv:2305.13606 simplifies this task by considering the soliton in a bounded region and relating the $\eta$ invariant, and thus the fermion number, to a specific heat kernel coefficient and to contributions from the edge states. We
On the detection of caustic crossing events associated with dark matter in the form of primordial black holes
astro-ph.COM. R. S. Hawkins
The possibility that stellar mass primordial black holes may make up at least a significant fraction of dark matter has recently received much attention, partly as a result of gravitational wave observations, but more specifically from observations of microlensing in the Galactic halo and in quasar gravitational lens systems. If this is the case then a numbe
Adam R. Klivans, Konstantinos Stavropoulos, Kevin Tian, Arsen Vasilyan
Inspired by recent work on learning with distribution shift, we give a general outlier removal algorithm called iterative polynomial filtering and show a number of striking applications for supervised learning with contamination: (1) We show that any function class that can be approximated by low-degree polynomials with respect to a hypercontractive distribu
Alkis Koudounas, Moreno La Quatra, Eliana Pastor, Sabato Marco Siniscalchi
Kolmogorov-Arnold Networks (KANs) have recently emerged as a promising alternative to traditional neural architectures, yet their application to speech processing remains under explored. This work presents the first investigation of KANs for Spoken Language Understanding (SLU) tasks. We experiment with 2D-CNN models on two datasets, integrating KAN layers in
URPlanner: A Universal Paradigm For Collision-Free Robotic Motion Planning Based on Deep Reinforcement Learning
cs.ROFengkang Ying, Hanwen Zhang, Haozhe Wang, Huishi Huang
Collision-free motion planning for redundant robot manipulators in complex environments is yet to be explored. Although recent advancements at the intersection of deep reinforcement learning (DRL) and robotics have highlighted its potential to handle versatile robotic tasks, current DRL-based collision-free motion planners for manipulators are highly costly,
Daryl. J. Daley, Yoni Nazarathy, Jiesen Wang
The paper studies the counting process arising as a subset of births and deaths in a birth--death process on a finite state space. Whenever a birth or death occurs, the process is incremented or not depending on the outcome of an independent Bernoulli experiment whose probability is a state-dependent function of the birth and death and also depends on whethe
Investigating Accretion Disk-Corona in Seyfert 1 galaxies: A UV/X-ray Spectral Study of Mrk 813 and RBS 688
astro-ph.HEPiyali Ganguly, Gulab C. Dewangan
We present a broadband UV/X-ray spectral study of two Seyfert 1 galaxies, Mrk 813 and RBS 688, primarily based on AstroSat observations. These active galactic nuclei host relatively large super-massive black holes ($M_{BH} \sim 10^8 - 10^9M_{\odot}$), suffer negligible internal extinction/absorption, and are well suited for probing the inner regions of their
Etienne Boursier, Scott Pesme, Radu-Alexandru Dragomir
We study the dynamics of gradient flow with small weight decay on general training losses $F: \mathbb{R}^d \to \mathbb{R}$. Under mild regularity assumptions and assuming convergence of the unregularised gradient flow, we show that the trajectory with weight decay $\lambda$ exhibits a two-phase behaviour as $\lambda \to 0$. During the initial fast phase, the
Ryan Po, Yotam Nitzan, Richard Zhang, Berlin Chen
Video diffusion models have recently shown promise for world modeling through autoregressive frame prediction conditioned on actions. However, they struggle to maintain long-term memory due to the high computational cost associated with processing extended sequences in attention layers. To overcome this limitation, we propose a novel architecture leveraging
Yunze Lin
Solving algebraic word problems (AWPs) has recently emerged as an important natural language processing task. Recently, large language models (LLMs) have demonstrated powerful mathematical capabilities, and the Chain-of-Thought technique, which guides LLMs through step-by-step reasoning, has yielded impressive results. However, this reasoning ability is limi
Structure and Elastic properties of Titanium MXenes: evaluation of COMB3, REAXFF and MEAM force fields
cond-mat.mtrl-sciLuis F. V. Thomazini, Alexandre F. Fonseca
Titanium carbide and nitride MXenes are two-dimensional inorganic materials that exhibit noteworthy physical and chemical properties. These materials are considered for a variety of technological applications, ranging from energy harvesting to optical and biomedical applications. Given the growing interest in titanium MXenes, there is an expanding demand for
Clément Berenfeld, Ahmed Boughdiri, Bénédicte Colnet, Wouter A. C. van Amsterdam
Meta-analysis, by synthesizing effect estimates from multiple studies conducted in diverse settings, stands at the top of the evidence hierarchy in clinical research. Yet, conventional approaches based on fixed- or random-effects models lack a causal framework, which may limit their interpretability and utility for public policy. Incorporating causal inferen
Leonel Bixano, Tonatiuh Matos
To date, many exotic predictions of Einstein's equations have been corroborated, with one exception: wormholes (WHs). In this work, we analyse a new exact solution combination to the Einstein-Maxwell-Dilaton or Phantom equations, which represent rotating WHs with magnetic and electric fields. The solution contains a ring singularity that, like other solution
Chun-Yi Kuan, Hung-yi Lee
Audio-aware large language models (ALLMs) have recently made great strides in understanding and processing audio inputs. These models are typically adapted from text-based large language models (LLMs) through additional training on audio-related tasks. This adaptation process presents two major limitations. First, ALLMs often suffer from catastrophic forgett
Tatiana Muraveva, Michele Bellazzini, Alessia Garofalo, Gisella Clementini
The Sagittarius (Sgr) dwarf spheroidal galaxy is one of the most prominent satellites of the Milky Way (MW). It is currently undergoing tidal disruption, forming an extensive stellar stream that provides key insights into the assembly history of the MW halo. In this study we analyzed RR Lyrae stars (RRLs) in the Sgr stream provided in Gaia Data Release 3 (DR
Dairu Liu, Ziyue Wang, Minyuan Ruan, Fuwen Luo
Images usually convey richer detail than text, but often include redundant information, which potentially downgrades multimodal reasoning performance. When faced with lengthy or complex messages, humans tend to employ abstract thinking to convert them into simple and concise abstracts. Inspired by this cognitive strategy, we introduce a novel paradigm to eli
Moreno La Quatra, Alkis Koudounas, Valerio Mario Salerno, Sabato Marco Siniscalchi
Despite the remarkable progress in end-to-end Automatic Speech Recognition (ASR) engines, accurately transcribing dysarthric speech remains a major challenge. In this work, we proposed a two-stage framework for the Speech Accessibility Project Challenge at INTERSPEECH 2025, which combines cutting-edge speech recognition models with LLM-based generative error
Alexander Panfilov, Paul Kassianik, Maksym Andriushchenko, Jonas Geiping
As large language models grow in capability and agency, identifying vulnerabilities through red-teaming becomes vital for safe deployment. However, traditional prompt-engineering approaches may prove ineffective once red-teaming turns into a \emph{weak-to-strong} problem, where target models surpass red-teamers in capabilities. To study this shift, we frame
Julián Tachella, Matthieu Terris, Samuel Hurault, Andrew Wang
DeepInverse is an open-source PyTorch-based library for solving imaging inverse problems. The library covers all crucial steps in image reconstruction from the efficient implementation of forward operators (e.g., optics, MRI, tomography), to the definition and resolution of variational problems and the design and training of advanced neural network architect
Many-body localization in a quantum Ising model with the long-range interaction: Accurate determination of the transition point
cond-mat.dis-nnIllia Lukin, Andrii Sotnikov, Alexander L. Burin
Many-body localization (MBL) transition emerges at strong disorder in interacting systems, separating chaotic and reversible dynamics. Although the existence of MBL transition within the macroscopic limit in spin chains with a short-range interaction was proved rigorously, the transition point is not found yet because of the dramatic sensitivity of the trans
Evaluating Software Plagiarism Detection in the Age of AI: Automated Obfuscation and Lessons for Academic Integrity
cs.SETimur Sağlam, Larissa Schmid
Plagiarism in programming assignments is a persistent issue in computer science education, increasingly complicated by the emergence of automated obfuscation attacks. While software plagiarism detectors are widely used to identify suspicious similarities at scale and are resilient to simple obfuscation techniques, they are vulnerable to advanced obfuscation
Patric Dolmeta, Matteo Giordano
We study nonparametric Bayesian inference for the intensity function of a covariate-driven point process. We extend recent results from the literature, showing that a wide class of Gaussian priors, combined with flexible link functions, achieve minimax optimal posterior contraction rates. Our result includes widespread prior choices such as the popular Mat\'
Yi Chen, Sen Liang, Zixiang Zhou, Ziyao Huang
Recent years have witnessed significant progress in audio-driven human animation. However, critical challenges remain in (i) generating highly dynamic videos while preserving character consistency, (ii) achieving precise emotion alignment between characters and audio, and (iii) enabling multi-character audio-driven animation. To address these challenges, we
Hanting Chen, Jiarui Qin, Jialong Guo, Tao Yuan
Large Language Models (LLMs) deliver state-of-the-art capabilities across numerous tasks, but their immense size and inference costs pose significant computational challenges for practical deployment. While structured pruning offers a promising avenue for model compression, existing methods often struggle with the detrimental effects of aggressive, simultane
UORA: Uniform Orthogonal Reinitialization Adaptation in Parameter-Efficient Fine-Tuning of Large Models
cs.CLXueyan Zhang, Jinman Zhao, Zhifei Yang, Yibo Zhong
This paper introduces Uniform Orthogonal Reinitialization Adaptation (UORA), a novel parameter-efficient fine-tuning (PEFT) approach for Large Language Models (LLMs). UORA achieves state-of-the-art performance and parameter efficiency by leveraging a low-rank approximation method to reduce the number of trainable parameters. Unlike existing methods such as L
The Harmonic Entropy Estimator: Minimax Optimality and Semiparametric Efficiency for Infinite Alphabets
math.STOctavio César Mesner
This paper considers the estimation of Shannon entropy for discrete distributions with countably infinite support. While minimax rates for finite-support distributions are established, infinite-support distributions present distinct challenges regarding bias control as probabilities vanish. We address this by introducing the \textit{harmonic entropy estimato
MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
cs.CVKai Sun, Yushi Bai, Zhen Yang, Jiajie Zhang
Large Multimodal Models (LMMs) typically build on ViTs (e.g., CLIP), yet their training with simple random in-batch negatives limits the ability to capture fine-grained visual differences, particularly in geometric scenarios. To address this challenge, we propose a novel hard negative contrastive learning framework for the vision encoder, which combines imag
The evolving categories multinomial distribution: introduction with applications to movement ecology and vote transfer
stat.APRicardo Carrizo Vergara, Marc Kéry, Trevor Hefley
We introduce the evolving categories multinomial (ECM) distribution for multivariate count data taken over time. This distribution models the counts of individuals following iid stochastic dynamics among categories, with the number and identity of the categories also evolving over time. We specify the one-time and two-times marginal distributions of the coun
Martin Guillemaud, Vera Dinkelacker, Mario Chavez
Multilayer networks offer a powerful framework for modeling complex systems across diverse domains, effectively capturing multiple types of connections and interdependent subsystems commonly found in real world scenarios. To analyze these networks, embedding techniques that project nodes into a lower-dimensional geometric space are essential. This paper intr
Ilai Reshef, Nadav Dym
Multiset functions, which are functions that map multisets to vectors, are a fundamental tool in the construction of neural networks for multisets and graphs. To guarantee that the vector representation of the multiset is faithful, it is often desirable to have multiset mappings that are both injective and bi-Lipschitz. Currently, there are several construct
Improvement Strategies for Few-Shot Learning in OCT Image Classification of Rare Retinal Diseases
eess.IVCheng-Yu Tai, Ching-Wen Chen, Chi-Chin Wu, Bo-Chen Chiu
This paper focuses on using few-shot learning to improve the accuracy of classifying OCT diagnosis images with major and rare classes. We used the GAN-based augmentation strategy as a baseline and introduced several novel methods to further enhance our model. The proposed strategy contains U-GAT-IT for improving the generative part and uses the data balance
Ziming Wei, Bingqian Lin, Zijian Jiao, Yunshuang Nie
Spatial Planning is a crucial part in the field of spatial intelligence, which requires the understanding and planning about object arrangements in space perspective. AI agents with the spatial planning ability can better adapt to various real-world applications, including robotic manipulation, automatic assembly, urban planning etc. Recent works have attemp
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
cs.CVJin Wang, Yao Lai, Aoxue Li, Shifeng Zhang
The rapid progress of large language models (LLMs) has catalyzed the emergence of multimodal large language models (MLLMs) that unify visual understanding and image generation within a single framework. However, most existing MLLMs rely on autoregressive (AR) architectures, which impose inherent limitations on future development, such as the raster-scan orde
Arthur S. de Sena, Jacek Kibilda, Nurul H. Mahmood, Andre Gomes
This article investigates the robustness of rate-splitting multiple access (RSMA) in multi-user multiple-input single-output (MISO) systems to interference attacks against channel acquisition induced by beyond-diagonal RISs (BD-RISs). Two primary attack strategies, random and aligned interference, are proposed for fully connected and group-connected reconfig
Auger parameter analysis for TiN and AlN thin films via combined in-situ XPS and HAXPES
cond-mat.mtrl-sciO. V. Pshyk, J. Patidar, C. Cancellieri, S. Siol
Auger parameter analysis provides in-depth information about the electronic and chemical bonding properties of TiN and AlN thin films, which are relevant across a wide range of technologies. Meaningful interpretation and analysis of the Auger parameter of these materials have been hindered due to, among other reasons, the absence of reliable references. Here
Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu Zhang
Despite the remarkable capabilities of Language Models (LMs) across diverse tasks, no single model consistently outperforms others, necessitating efficient methods to combine their strengths without expensive retraining. Existing model merging techniques, such as parameter averaging and task-guided fusion, often rely on data-dependent computations or fail to
Hanneke C. Woudenberg, Amina Helmi
Many of the Milky Way's accreted substructures have been discovered and studied in the space of energy $E$, and angular momentum components $L_z$ and $L_{\bot}$. In a static axisymmetric system, these quantities are (reasonable approximations of) the integrals of motion of an orbit. However, in a galaxy like the Milky Way with a triaxial, rotating bar, none
Sublimation of orientated amino acid films for reliable, amplified piezoelectric performance
physics.chem-phCiaran O Malley, Muhammad Usaid Memon, Krishna Hari, Tara Ryan
Biomolecular crystals, such as amino acids, peptides, and proteins, have emerged as potential next generation piezoelectric materials due to their low-cost, biocompatibility, eco-friendliness, and reduced permittivity versus ceramics. However, many challenges have limited their acceleration into mainstream sensing applications. Their natural self-assembly fr
Efe Yazgan, Pedro Silva
This Special Issue on "Top Quark at the New Physics Frontier" is devoted to the most massive fundamental elementary particle known, the top quark. The aim is to provide a comprehensive review of the current status and prospects of top quark physics at the Large Hadron Collider (LHC) and future colliders. We included articles that emphasize where the present
Jialin Yang, Dongfu Jiang, Lipeng He, Sherman Siu
As Large Language Models (LLMs) become integral to software development workflows, their ability to generate structured outputs has become critically important. We introduce StructEval, a comprehensive benchmark for evaluating LLMs' capabilities in producing both non-renderable (JSON, YAML, CSV) and renderable (HTML, React, SVG) structured formats. Unlike pr
FairTalk: Facilitating Balanced Participation in Video Conferencing by Implicit Visualization of Predicted Turn-Grabbing Intention
cs.HCRyo Iijima, Shigeo Yoshida, Atsushi Hashimoto, Jiaxin Ma
Creating fair opportunities for all participants to contribute is a notable challenge in video conferencing. This paper introduces FairTalk, a system that facilitates the subconscious redistribution of speaking opportunities. FairTalk predicts participants' turn-grabbing intentions using a machine learning model trained on web-collected videoconference data
Filippo Scaramuzza, Giovanni Quattrocchi, Damian A. Tamburri
As Artificial Intelligence (AI) systems, particularly those based on machine learning (ML), become integral to high-stakes applications, their probabilistic and opaque nature poses significant challenges to traditional verification and validation methods. These challenges are exacerbated in regulated sectors requiring tamper-proof, auditable evidence, as hig
Wenyang Liao, Quanziang Wang, Yichen Wu, Renzhen Wang
Replay-based continual learning (CL) methods assume that models trained on a small subset can also effectively minimize the empirical risk of the complete dataset. These methods maintain a memory buffer that stores a sampled subset of data from previous tasks to consolidate past knowledge. However, this assumption is not guaranteed in practice due to the lim
Lucrezia Bertoletti
Let $p$ be a prime number and $K$ a finite unramified extension of $\mathbb{Q}_p$. Building on recent work of Breuil, Herzig, Hu, Morra and Schraen, we study the smooth mod $p$ representations of $\mathrm{GL}_2(K)$ appearing in a tower of mod $p$ Hecke eigenspaces of the cohomology of Shimura curves, under mild genericity assumptions but notably no multiplic
Konstantin Dobler, Desmond Elliott, Gerard de Melo
Current language models rely on static vocabularies determined at pretraining time, which can lead to decreased performance and increased computational cost for domains underrepresented in the original vocabulary. New tokens can be added to solve this problem, when coupled with a good initialization for their new embeddings. However, existing embedding initi
Tensorization is a powerful but underexplored tool for compression and interpretability of neural networks
cs.LGSafa Hamreras, Sukhbinder Singh, Román Orús
Tensorizing a neural network involves reshaping some or all of its dense weight matrices into higher-order tensors and approximating them using low-rank tensor network decompositions. This technique has shown promise as a model compression strategy for large-scale neural networks. However, despite encouraging empirical results, tensorized neural networks (TN
What does making money have to do with crime?: A dive into the National Crime Victimization survey
physics.soc-phSydney Anuyah
In this short article, I leverage the National Crime Victimization Survey from 1992 to 2022 to examine how income, education, employment, and key demographic factors shape the type of crime victims experience (violent vs property). Using balanced classification splits and logistic regression models evaluated by F1-score, there is an isolation of the socioeco
MolEditRL: Structure-Preserving Molecular Editing via Discrete Diffusion and Reinforcement Learning
cs.LGYuanxin Zhuang, Dazhong Shen, Ying Sun
Molecular editing aims to modify a given molecule to optimize desired chemical properties while preserving structural similarity. However, current approaches typically rely on string-based or continuous representations, which fail to adequately capture the discrete, graph-structured nature of molecules, resulting in limited structural fidelity and poor contr
Balancing Interference and Correlation in Spatial Experimental Designs: A Causal Graph Cut Approach
cs.LGJin Zhu, Jingyi Li, Hongyi Zhou, Yinan Lin
This paper focuses on the design of spatial experiments to optimize the amount of information derived from the experimental data and enhance the accuracy of the resulting causal effect estimator. We propose a surrogate function for the mean squared error (MSE) of the estimator, which facilitates the use of classical graph cut algorithms to learn the optimal
Xinhang Liu, Yu-Wing Tai, Chi-Keung Tang
Despite recent advances in multimodal content generation enabled by vision-language models (VLMs), their ability to reason about and generate structured 3D scenes remains largely underexplored. This limitation constrains their utility in spatially grounded tasks such as embodied AI, immersive simulations, and interactive 3D applications. We introduce a new p
Zhengliang Shi, Lingyong Yan, Dawei Yin, Suzan Verberne
Large language models (LLMs) have been widely integrated into information retrieval to advance traditional techniques. However, effectively enabling LLMs to seek accurate knowledge in complex tasks remains a challenge due to the complexity of multi-hop queries as well as the irrelevant retrieved content. To address these limitations, we propose EXSEARCH, an
Fabiana Fournier, Lior Limonad, Yuval David
AI agents that leverage Large Language Models (LLMs) are increasingly becoming core building blocks of modern software systems. A wide range of frameworks is now available to support the specification of such applications. These frameworks enable the definition of agent setups using natural language prompting, which specifies the roles, goals, and tools assi
Shintaro Ito, Natsuki Takama, Toshiki Watanabe, Koichi Ito
Recent advancements in radiance field rendering, exemplified by Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), have significantly progressed 3D modeling and reconstruction. The use of multiple 360-degree omnidirectional images for these tasks is increasingly favored due to advantages in data acquisition and comprehensive scene capture. Howev
Parameshwar R. Pasnoori
In this letter we consider the time dependent Kondo model where a magnetic impurity interacts with the electrons through a time dependent interaction strength $J(t)$. We develop a new framework based on Bethe ansatz and construct an exact solution to the time-dependent Schrodinger equation. We show that when periodic boundary conditions are applied, the cons
Fanheng Kong, Jingyuan Zhang, Hongzhi Zhang, Shi Feng
Videos are unique in their integration of temporal elements, including camera, scene, action, and attribute, along with their dynamic relationships over time. However, existing benchmarks for video understanding often treat these properties separately or narrowly focus on specific aspects, overlooking the holistic nature of video content. To address this, we
Huijie Zhang, Zijian Huang, Siyi Chen, Jinfan Zhou
Diffusion distillation provides an effective approach for learning lightweight and few-steps diffusion models with efficient generation. However, evaluating their generalization remains challenging: theoretical metrics are often impractical for high-dimensional data, while no practical metrics rigorously measure generalization. In this work, we bridge this g
Anh Thai, Stefan Stojanov, Zixuan Huang, Bikram Boote
This paper introduces MEBench, a novel benchmark for evaluating mutual exclusivity (ME) bias, a cognitive phenomenon observed in children during word learning. Unlike traditional ME tasks, MEBench further incorporates spatial reasoning to create more challenging and realistic evaluation settings. To facilitate controlled experimentation, we also present a fl
Simpson Zhang, Tennison Liu, Mihaela van der Schaar
Current labor markets are strongly affected by the economic forces of adverse selection, moral hazard, and reputation, each of which arises due to $\textit{incomplete information}$. These economic forces will still be influential after AI agents are introduced, and thus, agents must use metacognitive and strategic reasoning to perform effectively. Metacognit
Jiaming Ma, Guanjun Wang, Sheng Huang, Kuo Yang
Due to the profound impact of air pollution on human health, livelihoods, and economic development, air quality forecasting is of paramount significance. Initially, we employ the causal graph method to scrutinize the constraints of existing research in comprehensively modeling the causal relationships between the air quality index (AQI) and meteorological fe
Dominik Meier, Jan Philip Wahle, Paul Röttger, Terry Ruas
As large language models (LLMs) become integrated into sensitive workflows, concerns grow over their potential to leak confidential information. We propose TrojanStego, a novel threat model in which an adversary fine-tunes an LLM to embed sensitive context information into natural-looking outputs via linguistic steganography, without requiring explicit contr
Joppe Stokvis
We study the 8-rank of class groups of hyperelliptic function fields and show that such 8-ranks are governed by splitting conditions in so-called governing fields. A similar result was proven for quadratic number fields by Stevenhagen, who used a theory of R\'edei symbols and R\'edei reciprocity to do so. We introduce a version of the R\'edei reciprocity law
The finite-difference parquet method: Enhanced electron-paramagnon scattering opens a pseudogap
cond-mat.str-elJae-Mo Lihm, Dominik Kiese, Seung-Sup B. Lee, Fabian B. Kugler
We present the finite-difference parquet method that greatly improves the applicability and accuracy of two-particle correlation approaches to interacting electron systems. This method incorporates the nonperturbative local physics from a reference solution and builds all parquet diagrams while circumventing potentially divergent irreducible vertices. Its un
Shadows and Observational Images of a Schwarzschild-like Black Hole Surrounded by a Dehnen-type Dark Matter Halo
gr-qcZuting Luo, Meirong Tang, Zhaoyi Xu
This paper investigates the optical appearance of a Schwarzschild-like black hole (BH) surrounded by a Dehnen-(1, 4, 5/2) type dark matter (DM) halo, with a focus on how the DM halo's density $\rho_{s}$ and radius $r_{s}$ influence the BH's shadow and photon ring. First, the radius $r_h$ of the BH's event horizon and the equation of motion for photons were d
Algorithmic Control Improves Residential Building Energy and EV Management when PV Capacity is High but Battery Capacity is Low
eess.SYLennart Ullner, Alona Zharova, Felix Creutzig
Efficient energy management in prosumer households is key to alleviating grid stress in an energy transition marked by electric vehicles (EV), renewable energies and battery storage. However, it is unclear how households optimize prosumer EV charging. Here we study real-world data from 90 households on fixed-rate electricity tariffs in German-speaking countr
Tri-band Aperture-shared Antenna Array Using Scalable FSS-based Electromagnetic Transparent Structure
physics.app-phYongzheng Li, Wanchen Yang, Quan Xue, Wenquan Che
In a tri-band aperture-shared array (TBA), the low-band (LB) dipole often deteriorates the radiation patterns of the middle-band (MB) and high-band (HB) antennas due to shielding effects. To address this issue, a novel dual-band electromagnetic transparent structure (DBTS) is firstly proposed and used to realize a TBA. The DBTS achieves two tunable electroma
Cristian Santini, Laura Melosi, Emanuele Frontoni
The increased digitization of world's textual heritage poses significant challenges for both computer science and literary studies. Overall, there is an urgent need of computational techniques able to adapt to the challenges of historical texts, such as orthographic and spelling variations, fragmentary structure and digitization errors. The rise of large lan
Haolei Bai, Siyong Jian, Tuo Liang, Yu Yin
Large language models (LLMs) have demonstrated impressive capabilities in a wide range of downstream natural language processing tasks. Nevertheless, their considerable sizes and memory demands hinder practical deployment, underscoring the importance of developing efficient compression strategies. Singular value decomposition (SVD) decomposes a matrix into o
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
cs.CLJian Lan, Udo Schlegel, Tanveer Hannan, Gengyuan Zhang
Although Vision-Language Models (VLMs) have achieved remarkable success, the knowledge mechanisms underlying their social biases remain a black box, where fairness- and ethics-related problems harm certain groups of people in society. It is unknown to what extent VLMs yield gender and race bias in generative responses. In this paper, we conduct a systematic
Kun Zhou, Zaiwu Gong, Guo Wei, Roman Slowinski
Limited by cognitive abilities, decision-makers (DMs) may struggle to evaluate decision alternatives based on all criteria in multiple criteria decision-making problems. This paper proposes an embedded criteria selection method derived from preference disaggregation technique and regularization theory. The method aims to infer the criteria and value function
June-Woo Kim, Wonkyo Oh, Haram Yoon, Sung-Hoon Yoon
Suicidal risk detection in adolescents is a critical challenge, yet existing methods rely on language-specific models, limiting scalability and generalization. This study introduces a novel language-agnostic framework for suicidal risk assessment with large language models (LLMs). We generate Chinese transcripts from speech using an ASR model and then employ
Classical-to-quantum transfer of geometric phase for non-interferometric phase measurement and manipulation of quantum state
quant-phVimlesh Kumar, Chahat Kaushik, M. Ebrahim-Zadeh, C. M. Chandrashekar
The geometric phase, originating from the cyclic evolution of a state, such as polarization on the Poincar\'e sphere, is typically measured through interferometric approaches that often include unwanted contributions from the dynamic phase. Here, we present a non-interferometric technique based on quantum correlation of pair photons to measure the geometric
Ziyi Zhang, Li Shen, Deheng Ye, Yong Luo
Text-to-multiview (T2MV) diffusion models have shown great promise in generating multiple views of a scene from a single text prompt. While few-step backbones enable real-time T2MV generation, they often compromise key aspects of generation quality, such as per-view fidelity and cross-view consistency. Reinforcement learning (RL) finetuning offers a potentia
Zuyao Chen, Jinlin Wu, Zhen Lei, Chang Wen Chen
We present OvSGTR, a novel transformer-based framework for fully open-vocabulary scene graph generation that overcomes the limitations of traditional closed-set models. Conventional methods restrict both object and relationship recognition to a fixed vocabulary, hindering their applicability to real-world scenarios where novel concepts frequently emerge. In
Realistic Multi-temperature Dust: How Well Can We Constrain the Dust Properties of High-redshift Galaxies?
astro-ph.GALaura Sommovigo, Hiddo Algera
Determining the dust properties of high-redshift galaxies from their far-infrared continuum emission is challenging due to limited multi-frequency data. As a result, the dust spectral energy distribution (SED) is often modeled as a single-temperature modified blackbody. We assess the accuracy of the single-temperature approximation by constructing realistic
Ivan Vybornyi, Shuying Chen, Lukas J. Spieß, Piet O. Schmidt
In quantum logic spectroscopy, internal transitions of trapped ions and molecules can be probed by measuring the motional displacement caused by an applied light field of variable frequency. This provides a solution to ``needle in a haystack'' problems, such as the search for narrow clock transitions in highly charged ions, recently discussed by S. Chen et a
Xiangyu Li, Jingqiang Chen
Citations are crucial in scientific research articles as they highlight the connection between the current study and prior work. However, this process is often time-consuming for researchers. In this study, we propose the SciRGC framework, which aims to automatically recommend citation articles and generate citation sentences for citation locations within ar
On a family of continued fractions in $Q((T^1))$ associated to infinite binary words derived from the Thue-Morse sequence
math.NTBill Allombert, Alain Lasjaunias
For each integer n > 1, we present an element in $Q((T^-1))$, having a power series expansion based on an infinite word W(n), over the alphabet ${+1;-1}g and whose continued fraction expansion has a particular pattern which is explicitly described. The word W(1) is the Thue-Morse sequence and the following words are defined in a similar way.
Yunhao Wang, Yuhao Zhang, Tinghao Yu, Can Xu
Large language models (LLMs) have shown impressive capabilities in handling complex tasks through long-chain reasoning. However, the extensive reasoning steps involved can significantly increase computational costs, posing challenges for real-world deployment. Recent efforts have focused on optimizing reasoning efficiency by shortening the Chain-of-Thought (
Fengyuan Sun, Leqi Shen, Hui Chen, Sicheng Zhao
Video Large Language Models (Video LLMs) have achieved remarkable results in video understanding tasks. However, they often suffer from heavy computational overhead due to the large number of visual tokens generated from multiple video frames. Existing visual token compression methods often rely on attention scores from language models as guidance. However,
Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and Opportunities
cs.CLChuangtao Ma, Yongrui Chen, Tianxing Wu, Arijit Khan
Large language models (LLMs) have demonstrated remarkable performance on question-answering (QA) tasks because of their superior capabilities in natural language understanding and generation. However, LLM-based QA struggles with complex QA tasks due to poor reasoning capacity, outdated knowledge, and hallucinations. Several recent works synthesize LLMs and k
Xiaowen Ling, Zhiqiang Li, Yanbin Wang, Zhuhong You
As protein informatics advances rapidly, the demand for enhanced predictive accuracy, structural analysis, and functional understanding has intensified. Transformer models, as powerful deep learning architectures, have demonstrated unprecedented potential in addressing diverse challenges across protein research. However, a comprehensive review of Transformer
Liang Cheng, Tianyi LI, Zhaowei Wang, Mark Steedman
The performance of pre-trained Large Language Models (LLMs) is often sensitive to nuances in prompt templates, requiring careful prompt engineering, adding costs in terms of computing and human effort. In this study, we present experiments encompassing multiple LLMs variants of varying sizes aimed at probing their preference with different prompts. Through e
MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
cs.CLThang Nguyen, Peter Chin, Yu-Wing Tai
We present MA-RAG, a Multi-Agent framework for Retrieval-Augmented Generation (RAG) that addresses the inherent ambiguities and reasoning challenges in complex information-seeking tasks. Unlike conventional RAG methods that rely on end-to-end fine-tuning or isolated component enhancements, MA-RAG orchestrates a collaborative set of specialized AI agents: Pla
Chenxiang Zhang, Jun Pang, Sjouke Mauw
Neural networks trained on real-world data often exhibit biases while simultaneously being vulnerable to privacy attacks aimed at extracting sensitive information. Despite extensive research on each problem individually, their intersection remains poorly understood. In this work, we investigate the privacy impact of spurious correlation bias. We introduce \e
Qi Li, Kun Li, Haozhi Han, Honghui Shang
Can a scientific simulation system be physically consistent, interpretable by design, and scalable across regimes--all at once? Despite decades of progress, this trifecta remains elusive. Classical methods like Kinetic Monte Carlo ensure thermodynamic accuracy but scale poorly; learning-based methods offer efficiency but often sacrifice physical consistency
Technical recommendation on multiplex MR elastography for tomographic mapping of abdominal stiffness with a focus on the pancreas and pancreatic ductal adenocarcinoma
physics.med-phJakob Schattenfroh, Salma Almutawakel, Jan Bieling, Johannes Castelein
Objectives: MR elastography (MRE) offers valuable mechanical tissue characterization, however, in deep abdominal organs like the pancreas conventional single-driver, single-frequency approaches often fail. This study evaluates whether multiplex MRE using multiple drivers and vibration frequencies can overcome these limitations. Methods: This study used singl