May 2025 arXiv papers — page 66
Showing 6,501–6,600 of 24,552 papers
Moderating Harm: Benchmarking Large Language Models for Cyberbullying Detection in YouTube Comments
cs.CLAmel Muminovic
As online platforms grow, comment sections increasingly host harassment that undermines user experience and well-being. This study benchmarks three leading large language models, OpenAI GPT-4.1, Google Gemini 1.5 Pro, and Anthropic Claude 3 Opus, on a corpus of 5,080 YouTube comments sampled from high-abuse threads in gaming, lifestyle, food vlog, and music
Jingxuan Xu, Hong Huang, Chuhang Zou, Manolis Savva
We propose a neural physics system for real-time, interactive fluid simulations. Traditional physics-based methods, while accurate, are computationally intensive and suffer from latency issues. Recent machine-learning methods reduce computational costs while preserving fidelity; yet most still fail to satisfy the latency constraints for real-time use and lac
Words as Geometric Features: Estimating Homography using Optical Character Recognition as Compressed Image Representation
cs.CVRoss Greer, Alisha Ukani, Katherine Izhikevich, Earlence Fernandes
Document alignment and registration play a crucial role in numerous real-world applications, such as automated form processing, anomaly detection, and workflow automation. Traditional methods for document alignment rely on image-based features like keypoints, edges, and textures to estimate geometric transformations, such as homographies. However, these appr
Chenxi Li, Nuo Chen, Fengyun Tan, Yantong Chen
We present a novel active learning framework for 3D point cloud semantic segmentation that, for the first time, integrates large language models (LLMs) to construct hierarchical label structures and guide uncertainty-based sample selection. Unlike prior methods that treat labels as flat and independent, our approach leverages LLM prompting to automatically g
Yile Li, Shandian Zhe
Operator learning seeks to approximate mappings from input functions to output solutions, particularly in the context of partial differential equations (PDEs). While recent advances such as DeepONet and Fourier Neural Operator (FNO) have demonstrated strong performance, they often rely on regular grid discretizations, limiting their applicability to complex
Ab Initio Prediction of Large Thermoelectric Effect in Distorted Heusler Alloy Ti-Fe-Sb Compound
cond-mat.mtrl-sciRifky Syariati, Athorn Vora-ud, Fumiyuki Ishii, Tosawat Seetawan
The thermoelectric figure of merit of the Heusler alloy TiFe$_{1.5}$Sb was investigated by first-principles calculations of lattice thermal conductivity. The electronic thermal conductivity, electrical conductivity, and Seebeck coefficient are calculated by semi-classical Boltzmann transport theory. TiFe$_{1.5}$Sb was found to be thermally and dynamically st
Nima Shahbazi, Stavros Sintos, Abolfazl Asudeh
Frequency estimation in streaming data often relies on sketches like Count-Min (CM) to provide approximate answers with sublinear space. However, CM sketches introduce additive errors that disproportionately impact low-frequency elements, creating fairness concerns across different groups of elements. We introduce Fair-Count-Min, a frequency estimation sketc
Javier Salazar Cavazos, Jeffrey A Fessler, Laura Balzano
Principal component analysis (PCA) is a key tool in the field of data dimensionality reduction. Various methods have been proposed to extend PCA to the union of subspace (UoS) setting for clustering data that comes from multiple subspaces like K-Subspaces (KSS). However, some applications involve heterogeneous data that vary in quality due to noise character
Zhepeng Cen, Yihang Yao, William Han, Zuxin Liu
Reinforcement learning (RL) has emerged as a powerful post-training technique to incentivize the reasoning ability of large language models (LLMs). However, LLMs can respond very inconsistently to RL finetuning: some show substantial performance gains, while others plateau or even degrade. To understand this divergence, we analyze the per-step influence of t
Yue Li, Jake Vasilakes, Zhixue Zhao, Carolina Scarton
We introduce SCRum-9, the largest multilingual Stance Classification dataset for Rumour analysis in 9 languages, containing 7,516 tweets from X. SCRum-9 goes beyond existing stance classification datasets by covering more languages, linking examples to more fact-checked claims (2.1k), and including confidence-related annotations from multiple annotators to a
Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering
cs.CVYixiong Chen, Wenjie Xiao, Pedro R. A. S. Bassi, Xinze Zhou
Vision-Language Models (VLMs) have shown promise in various 2D visual tasks, yet their readiness for 3D clinical diagnosis remains unclear due to stringent demands for recognition precision, reasoning ability, and domain knowledge. To systematically evaluate these dimensions, we present DeepTumorVQA, a diagnostic visual question answering (VQA) benchmark tar
Carlo Musolino
In this paper, we present an approach to solving the Riemann problem in one-dimensional relativistic hydrodynamics, where the most computationally expensive steps of the exact solver are replaced by compact, highly specialized neural networks. The resulting "neural" Riemann solver is integrated into a high-resolution shock-capturing scheme and tested on a ra
Surajit Sen, Tushar Kanti Dey, Anushree Bhattacharjee, Sovik Roy
We demonstrate quantum teleportation of a qutrit system using a complete set of two-qutrit entangled states obtained from the representation theory of the SU(3) group. All measurement gates essential for end-to-end teleportation are systematically evaluated, and these are found to be non-unitary. Our approach extends Bennett's teleportation protocol to t
Hamidreza Montazeri Hedesh, Moh. Kamalul Wafi, Bahram Shafai, Milad Siami
This paper investigates the robustness of the Lur'e problem under positivity constraints, drawing on results from the positive Aizerman conjecture and robustness properties of Metzler matrices. Specifically, we consider a control system of Lur'e type in which not only the linear part includes parametric uncertainty but also the nonlinear sector bound is unkn
Debayan Das, Antonio Cutrona, Andrew C. Cooper, Luana Olivieri
Microcombs require ultralow-noise repetition rates to enable next-generation applications in metrology, high-speed communications, microwave photonics, and sensing, where spectral purity is a central performance metric. Best-performing sources operate actively locked at "quiet points" in parameter space, fixed by device and material properties. Creating broa
Uncovering relationships between the electronic self-energy and coupled-cluster doubles theory
cond-mat.str-elChristopher J. N. Coveney
We derive the coupled-cluster doubles (CCD) amplitude equations by introduction of the particle-hole-time decoupled electronic self-energy. The resulting analysis leads to an expression for the ground state correlation energy that is exactly of the form obtained in coupled-cluster doubles theory. We demonstrate the relationship to the ionization potential/el
Andi Han, Wei Huang, Zhanpeng Zhou, Gang Niu
Deep learning with noisy labels presents significant challenges. In this work, we theoretically characterize the role of label noise from a feature learning perspective. Specifically, we consider a signal-noise data distribution, where each sample comprises a label-dependent signal and label-independent noise, and rigorously analyze the training dynamics of
Long-Duration Nonthermal Motions in the Supra-Arcade and Loop-Top Region During an Eruptive Solar Flare
astro-ph.SRTingyu Gou, Katharine K. Reeves
Solar flares are widely accepted to be powered by magnetic reconnection that involves complex dynamics in various scales. The flare supra-arcade and loop-top region, directly impacted by fast reconnection downflows, contains a wealth of microscopic dynamics, which are, however, difficult to resolve in imaging. We present simultaneous spectroscopic and imagin
Sanjay Kariyappa, G. Edward Suh
Prompt injection attacks are a critical security vulnerability in large language models (LLMs), allowing attackers to hijack model behavior by injecting malicious instructions within the input context. Recent defense mechanisms have leveraged an Instruction Hierarchy (IH) Signal, often implemented through special delimiter tokens or additive embeddings to de
Jonathan Hellwig, André Platzer
This paper introduces a proof calculus for real-analytic differential-algebraic dynamic logic, enabling correct transformations of differential-algebraic equations. Applications include index reductions from differential-algebraic equations to ordinary differential equations. The calculus ensures compatibility between differential-algebraic equation proof pr
Annalena Kofler, Vincent Stimper, Mikhail Mikhasenko, Michael Kagan
High-energy physics requires the generation of large numbers of simulated data samples from complex but analytically tractable distributions called matrix elements. Surrogate models, such as normalizing flows, are gaining popularity for this task due to their computational efficiency. We adopt an approach based on Flow Annealed importance sampling Bootstrap
Abhijit Chakraborty, Chahana Dahal, Vivek Gupta
Federated Retrieval-Augmented Generation (Federated RAG) combines Federated Learning (FL), which enables distributed model training without exposing raw data, with Retrieval-Augmented Generation (RAG), which improves the factual accuracy of language models by grounding outputs in external knowledge. As large language models are increasingly deployed in priva
Jakob Deser, Raphael Römer, Niklas Boers, Christian Kuehn
In bistable dynamical systems driven by Wiener processes, the widely used Kramers' law relates the strength of the noise forcing to the average time it takes to see a noise-induced transition from one attractor to the other. We extend this law to bistable systems forced by fast chaotic dynamics, which we argue is in some cases a more realistic modeling appro
Valentin Barriere, Nahuel Gomez, Leo Hemamou, Sofia Callejas
Aiming towards improving current computational models of humor detection, we propose a new multimodal dataset of stand-up comedies, in seven languages: English, French, Spanish, Italian, Portuguese, Hungarian and Czech. Our dataset of more than 330 hours, is at the time of writing the biggest available for this type of task, and the most diverse. The whole d
Laura Baracaldo, Blythe King, Haoran Yan, Yizi Lin
Cell boundary information is crucial for analyzing cell behaviors from time-lapse microscopy videos. Existing supervised cell segmentation tools, such as ImageJ, require tuning various parameters and rely on restrictive assumptions about the shape of the objects. While recent supervised segmentation tools based on convolutional neural networks enhance accura
Xiaoyan Hu, Lauren Pick, Ho-fung Leung, Farzan Farnia
The rapid advancement of generative AI has provided users with a wide range of well-trained models to address diverse prompts. When selecting a model for a given prompt, users should weigh not only its performance but also its service cost. However, existing model-selection methods typically emphasize performance while overlooking cost differences. In this p
J. W. Moffat, E. J. Thompson
We investigate the dynamical properties of dark energy through a detailed analysis of its equation of state parameter $w(z)$ as a function of redshift. We derive a general expression for $w(z)$ from the Friedmann-Lema\^itre-Robertson-Walker (FLRW) equations, establishing a direct relationship between the dark energy equation of state and the observable Hubbl
Beyond Domain Randomization: Event-Inspired Perception for Visually Robust Adversarial Imitation from Videos
cs.CVAndrea Ramazzina, Vittorio Giammarino, Matteo El-Hariry, Mario Bijelic
Imitation from videos often fails when expert demonstrations and learner environments exhibit domain shifts, such as discrepancies in lighting, color, or texture. While visual randomization partially addresses this problem by augmenting training data, it remains computationally intensive and inherently reactive, struggling with unseen scenarios. We propose a
Absence of Long-Range Order and Magnetic Anisotropy in the Triangular Magnet NdMgAl$_{11}$O$_{19}$
cond-mat.str-elSonu Kumar, Gaël Bastien, Jan Prokleška, Mateusz Kempiński
We investigated the rare-earth triangular-lattice antiferromagnet NdMgAl$_{11}$O$_{19}$ using single-crystal magnetization (1.8~K $\leq T \leq$ 300~K, $\mu_0 H \leq 7$~T) and specific-heat measurements down to 45~mK. The dc susceptibility confirms a well-isolated Kramers doublet ground state with pronounced Ising-type anisotropy, with $g_c \approx 3.7$ and $
Sonal Prabhune, Balaji Padmanabhan, Kaushik Dutta
We investigate the existence and persistence of a specific type of gender bias in some of the popular LLMs and contribute a new benchmark dataset, RealWorldQuestioning (released on HuggingFace ), developed from real-world questions across four key domains in business and health contexts: education, jobs, personal financial management, and general health. We
Dipanwita Saha, Anis Zaman, Hua Zou, Ning Chen
In search advertising, keyword matching connects user queries with relevant ads. While token-based matching increases ad coverage, it can reduce relevance due to overly permissive semantic expansion. This work extends keyword reach through document-side semantic keyword expansion, using a language model to broaden token-level matching without altering querie
Johannes Hofscheier, Vadym Kurylenko, Benjamin Nill
Lattice polytopes are called IDP polytopes if they have the integer decomposition property, i.e., any lattice point in a $k$th dilation is a sum of $k$ lattice points in the polytope. It is a long-standing conjecture whether the numerator of the Ehrhart series of an IDP polytope, called the $h^*$-polynomial, has a unimodal coefficient vector. In this prelimi
Fei Huang, Silvana M. Pesenti
This paper introduces marginal fairness, a new individual fairness notion for equitable decision-making in the presence of protected attributes such as gender, race, and religion. This criterion ensures that decisions based on generalized distortion risk measures are insensitive to distributional perturbations in protected attributes, regardless of whether t
Vanessa Utz, Steve DiPaola
Generative Artificial Intelligence (AI) systems currently contribute negatively to the production of digital waste, via the associated energy consumption and the related CO2 emissions. At this moment, a discussion is urgently needed on the replication of harmful consumer behavior, such as overconsumption, in the digital space. We outline our previous work on
Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
cs.CVWeixing Wang, Zifeng Ding, Jindong Gu, Rui Cao
Large Vision-Language Models (LVLMs) with discrete image tokenizers unify multimodal representations by encoding visual inputs into a finite set of tokens. Despite their effectiveness, we find that these models still hallucinate non-existent objects. We hypothesize that this may be due to visual priors induced during training: When certain image tokens frequ
Reva Schwartz, Rumman Chowdhury, Akash Kundu, Heather Frase
Conventional AI evaluation approaches concentrated within the AI stack exhibit systemic limitations for exploring, navigating and resolving the human and societal factors that play out in real world deployment such as in education, finance, healthcare, and employment sectors. AI capability evaluations can capture detail about first-order effects, such as whe
Vanessa Utz, Steve DiPaola
Climate implications of rapidly developing digital technologies, such as blockchains and the associated crypto mining and NFT minting, have been well documented and their massive GPU energy use has been identified as a cause for concern. However, we postulate that due to their more mainstream consumer appeal, the GPU use of text-prompt based diffusion AI art
Degradation-Aware and Machine Learning-Driven Uncertainty Quantification in Crystal Plasticity Finite Element: Texture-Driven Plasticity in 316L Stainless Steel
stat.APDinesh Kumar, Eralp Demir, Julio Spadotto, Kazuma Kobayashi
The mechanical properties and long-term structural reliability of crystalline materials are strongly influenced by microstructural features such as grain size, morphology, and crystallographic texture. These characteristics not only determine the initial mechanical behavior but also govern the progression of degradation mechanisms, such as strain localizatio
Morteza Rakhshaninejad, Mira Jurgens, Nicolas Dewolf, Willem Waegeman
Accurate drug-target interaction (DTI) prediction with machine learning models is essential for drug discovery. Such models should also provide a credible representation of their uncertainty, but applying classical marginal conformal prediction (CP) in DTI prediction often overlooks variability across drug and protein subgroups. In this work, we analyze thre
Miles Q. Li, Benjamin C. M. Fung
Large Language Models (LLMs) such as ChatGPT and its competitors have caused a revolution in natural language processing, but their capabilities also introduce new security vulnerabilities. This survey provides a comprehensive overview of these emerging concerns, categorizing threats into several key areas: inference-time attacks via prompt manipulation; tra
Deepak Kumar, Tuhin Malik, Hiranmaya Mishra, Constança Providência
A proto-neutron star (PNS) gets formed after a successful supernova when the stellar remnant decouples from the ejecta. In this study, we explore a relativistic framework for the finite-temperature $\beta$-equilibrium limit of equation of state (EOS), constrained via a Bayesian inference methodology. The EOS is constrained by minimal approximations on a few
Temperature- and charge carrier density-dependent electronic response in methylammonium lead iodide
cond-mat.mtrl-sciJiacheng Wang Jungmin Park, Lei Gao, Lucia Di Virgilio, Sheng Qu
Understanding carrier dynamics in photoexcited metal-halide perovskites is key for optoelectronic devices such as solar cells (low carrier densities) and lasers (high carrier densities). Trapping processes at low carrier densities and many-body recombination at high densities can significantly alter the dynamics of photoexcited carriers. Combining optical-pu
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
cs.LGZhendong Mi, Qitao Tan, Xiaodong Yu, Zining Zhu
Large language models (LLMs) have demonstrated impressive capabilities across numerous NLP tasks. Nevertheless, conventional first-order fine-tuning techniques impose heavy memory demands, creating practical obstacles to real-world applications. Zeroth-order (ZO) optimization has recently emerged as a promising memory-efficient alternative, as it circumvents
Alexander Erhardt, Alexander Wolff
The \emph{linear vertex arboricity} of a graph is the smallest number of sets into which the vertices of a graph can be partitioned so that each of these sets induces a linear forest. Chaplick et al. [JoCG 2020] showed that, somewhat surprisingly, the linear vertex arboricity of a graph is the same as the \emph{3D weak line cover number} of the graph, that i
Borna Khodabandeh, Amirabbas Afzali, Amirhossein Afsharrad, Seyed Shahabeddin Mousavi
Visual encoders have become fundamental components in modern computer vision pipelines. However, ensuring robustness against adversarial perturbations remains a critical challenge. Recent efforts have explored both supervised and unsupervised adversarial fine-tuning strategies. We identify two key limitations in these approaches: (i) they often suffer from i
Justin Deschenaux, Lan Tran, Caglar Gulcehre
Masked generative models (MGMs) can generate tokens in parallel and in any order, unlike autoregressive models (ARMs), which decode one token at a time, left-to-right. However, MGMs process the full-length sequence at every sampling step, including mask tokens that carry no information. In contrast, ARMs process only the previously generated tokens. We intro
Yuchen Wu, Edward Sun, Kaijie Zhu, Jianxun Lian
Large language models (LLMs) typically generate identical or similar responses for all users given the same prompt, posing serious safety risks in high-stakes applications where user vulnerabilities differ widely. Existing safety evaluations primarily rely on context-independent metrics - such as factuality, bias, or toxicity - overlooking the fact that the
SD-OVON: A Semantics-aware Dataset and Benchmark Generation Pipeline for Open-Vocabulary Object Navigation in Dynamic Scenes
cs.CVDicong Qiu, Jiadi You, Zeying Gong, Ronghe Qiu
We present the Semantics-aware Dataset and Benchmark Generation Pipeline for Open-vocabulary Object Navigation in Dynamic Scenes (SD-OVON). It utilizes pretraining multimodal foundation models to generate infinite unique photo-realistic scene variants that adhere to real-world semantics and daily commonsense for the training and the evaluation of navigation
Weihan Xu, Yimeng Ma, Jingyue Huang, Yang Li
Short videos are an effective tool for promoting contents and improving knowledge accessibility. While existing extractive video summarization methods struggle to produce a coherent narrative, existing abstractive methods cannot `quote' from the input videos, i.e., inserting short video clips in their outputs. In this work, we explore novel video editing mod
Thomas L. Draper, Feras A. Saad
This article studies the fundamental problem of using i.i.d. coin tosses from an entropy source to efficiently generate random variables $X_i \sim P_i$ $(i \ge 1)$, where $(P_1, P_2, \dots)$ is a random sequence of rational discrete probability distributions subject to an \textit{arbitrary} stochastic process. Our method achieves an amortized expected entrop
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
cs.CLKung-Hsiang Huang, Akshara Prabhakar, Onkar Thorat, Divyansh Agarwal
While AI agents hold transformative potential in business, effective performance benchmarking is hindered by the scarcity of public, realistic business data on widely used platforms. Existing benchmarks often lack fidelity in their environments, data, and agent-user interactions, with limited coverage of diverse business scenarios and industries. To address
DiffusionRL: Efficient Training of Diffusion Policies for Robotic Grasping Using RL-Adapted Large-Scale Datasets
cs.ROMaria Makarova, Qian Liu, Dzmitry Tsetserukou
Diffusion models have been successfully applied in areas such as image, video, and audio generation. Recent works show their promise for sequential decision-making and dexterous manipulation, leveraging their ability to model complex action distributions. However, challenges persist due to the data limitations and scenario-specific adaptation needs. In this
Sajal Chakroborty, Suddhasattwa Das
All techniques for denoising involve a notion of a true (noise-free) image, and a hypothesis space. The hypothesis space may reconstruct the image directly as a grayscale valued function, or indirectly by its Fourier or wavelet spectrum. Most common techniques estimate the true image as a projection to some subspace. We propose an interpretation of a noisy i
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
cs.CVShuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li
Diffusion Transformers (DiTs) are essential for video generation but suffer from significant latency due to the quadratic complexity of attention. By computing only critical tokens, sparse attention reduces computational costs and offers a promising acceleration approach. However, we identify that existing methods fail to approach optimal generation quality
Yiyu Ni, Marine A. Denolle, Amanda M. Thomas, Alex Hamilton
We present the first global-scale database of 4.3 billion P- and S-wave picks extracted from 1.3 PB continuous seismic data via a cloud-native workflow. Using cloud computing services on Amazon Web Services, we launched ~145,000 containerized jobs on continuous records from 47,354 stations spanning 2002-2025, completing in under three days. Phase arrivals we
Caitlin M. Casey, Hollis B. Akins, Steven L. Finkelstein, Maximilien Franco
By virtue of their red color, the dust in little red dots (LRDs) has been thought to be of appreciable influence, whether that dust is distributed in a torus around a compact active galactic nucleus (AGN) or diffuse in the interstellar medium (ISM) of nascent galaxies. In Casey et al. (2024) we predicted that, based on the compact sizes of LRDs (unresolved i
Brady D. Lund, Tae Hee Lee, Ziang Wang, Ting Wang
In response to the increasing complexity and sophistication of cyber threats, particularly those enhanced by advancements in artificial intelligence, traditional security methods are proving insufficient. This paper explores the Zero Trust cybersecurity framework, which operates on the principle of never trust, always verify to mitigate vulnerabilities withi
Nicolas Nguyen, Solenne Gaucher, Claire Vernade
We study the problem of non-stationary Lipschitz bandits, where the number of actions is infinite and the reward function, satisfying a Lipschitz assumption, can change arbitrarily over time. We design an algorithm that adaptively tracks the recently introduced notion of significant shifts, defined by large deviations of the cumulative reward function. To de
Yuri Miyagi, Nils Rodrigues, Daniel Weiskopf, Takayuki Itoh
By analyzing the gaze trajectories of people viewing screens and advertisements, we can determine what people are interested in. This knowledge can be effective when recommending commercial products and services, and also, when improving advertisement design. Therefore, analysis and visualization of eye gaze have been an active research topic. This paper pro
Understanding the Relationship Between Personal Data Privacy Literacy and Data Privacy Information Sharing by University Students
cs.CRBrady D. Lund, Bryan Anderson, Ana Roeschley, Gahangir Hossain
With constant threats to the safety of personal data in the United States, privacy literacy has become an increasingly important competency among university students, one that ties intimately to the information sharing behavior of these students. This survey based study examines how university students in the United States perceive personal data privacy and
Ankan Dash, Jingyi Gu, Guiling Wang, Chen Chen
Virtual Reality (VR) headsets, while integral to the evolving digital ecosystem, present a critical challenge: the occlusion of users' eyes and portions of their faces, which hinders visual communication and may contribute to social isolation. To address this, we introduce RevAvatar, an innovative framework that leverages AI methodologies to enable reverse p
Emilio de Carvalho, Percy Fernández-Sánchez, Marcelo Escudeiro Hernandes
We present some results concerning the Saito module and the torsion submodule of an analytic plane curve, and we provide a method for computing them. Using this algorithm, we compute analytic invariants for plane curves with multiplicity less than or equal to three.
Ming Cheng, Jiaying Gong, Hoda Eldardiry
Lay paraphrasing aims to make scientific information accessible to audiences without technical backgrounds. However, most existing studies focus on a single domain, such as biomedicine. With the rise of interdisciplinary research, it is increasingly necessary to comprehend knowledge spanning multiple technical fields. To address this, we propose Sci-LoRA, a
Md Farhamdur Reza, Reza Jahani, Richeng Jin, Huaiyu Dai
Decentralized federated learning (DFL) has attracted significant attention due to its scalability and independence from a central server. In practice, some participating clients can be mobile, yet the impact of user mobility on DFL performance remains largely unexplored, despite its potential to facilitate communication and model convergence. In this work, w
SW-ViT: A Spatio-Temporal Vision Transformer Network with Post Denoiser for Sequential Multi-Push Ultrasound Shear Wave Elastography
eess.IVAhsan Habib Akash, MD Jahin Alam, Md. Kamrul Hasan
Objective: Ultrasound Shear Wave Elastography (SWE) demonstrates great potential in assessing soft-tissue pathology by mapping tissue stiffness, which is linked to malignancy. Traditional SWE methods have shown promise in estimating tissue elasticity, yet their susceptibility to noise interference, reliance on limited training data, and inability to generate
Binhao Ma, Hanqing Guo, Zhengping Jay Luo, Rui Duan
Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced the naturalness and flexibility of human computer interaction by enabling seamless understanding across text, vision, and audio modalities. Among these, voice enabled models such as SpeechGPT have demonstrated considerable improvements in usability, offering expressive, a
Stanislav Semenov
We introduce and investigate the concept of Stratified Algebra, a new algebraic framework equipped with a layer-based structure on a vector space. We formalize a set of axioms governing intra-layer and inter-layer interactions, study their implications for algebraic dynamics, and present concrete matrix-based models that satisfy different subsets of these ax
Subek Acharya, Sansrit Paudel
Parkinson's Disease (PD) is a neurodegenerative disorder that significantly impacts motor and non-motor functions. There is currently no treatment that slows or stops neurodegeneration in PD. In this context, assistive technologies (ATs) have emerged as vital tools to aid people with Parkinson's and significantly improve their quality of life. This review ex
Corentin Delacour, M Mahmudul Hasan Sajeeb, Joao P. Hespanha, Kerem Y. Camsari
Sampling Boltzmann probability distributions plays a key role in machine learning and optimization, motivating the design of hardware accelerators such as Ising machines. While the Ising model can in principle encode arbitrary optimization problems, practical implementations are often hindered by soft constraints that either slow down mixing when too strong,
Securing Credit Inquiries: The Role of Real-Time User Approval in Preventing SSN Identity Theft
cs.CRGogulakrishnan Thiyagarajan, Vinay Bist, Prabhudarshi Nayak
Unauthorized credit inquiries are also a central entry point for identity theft, with Social Security Numbers (SSNs) being widely utilized in fraudulent cases. Traditional credit inquiry systems do not usually possess strict user authentication, making them vulnerable to unauthorized access. This paper proposes a real-time user authorization system to enhanc
Masao Someki, Shikhar Bharadwaj, Atharva Anand Joshi, Chyi-Jiunn Lin
Speech foundation models achieve strong generalization across languages and acoustic conditions, but require significant computational resources for inference. In the context of speech foundation models, pruning techniques have been studied that dynamically optimize model structures based on the target audio leveraging external context. In this work, we exte
Yuxiang Liu, Kevin Chen-Chuan Chang
We introduce the Exemplar-Based Expository Text Generation task, aiming to generate an expository text on a new topic using an exemplar on a similar topic. Current methods fall short due to their reliance on extensive exemplar data, difficulty in adapting topic-specific content, and issues with long-text coherence. To address these challenges, we propose the
Making deep neural networks work for medical audio: representation, compression and domain adaptation
cs.SDCharles C Onu
This thesis addresses the technical challenges of applying machine learning to understand and interpret medical audio signals. The sounds of our lungs, heart, and voice convey vital information about our health. Yet, in contemporary medicine, these sounds are primarily analyzed through auditory interpretation by experts using devices like stethoscopes. Autom
Maeva Guerrier, Karthik Soma, Hassan Fouad, Giovanni Beltrame
Safety stands as the primary obstacle preventing the widespread adoption of learning-based robotic systems in our daily lives. While reinforcement learning (RL) shows promise as an effective robot learning paradigm, conventional RL frameworks often model safety by using single scalar negative rewards with immediate episode termination, failing to capture the
Hierarchical-embedding autoencoder with a predictor (HEAP) as efficient architecture for learning long-term evolution of complex multi-scale physical systems
cs.AIAlexander Khrabry, Edward Startsev, Andrew Powis, Igor Kaganovich
We propose a novel efficient architecture for learning long-term evolution in complex multi-scale physical systems which is based on the idea of separation of scales. Structures of various scales that dynamically emerge in the system interact with each other only locally. Structures of similar scale can interact directly when they are in contact and indirect
Chika Maduabuchi, Hao Chen, Yujin Han, Jindong Wang
Latent Video Diffusion Models (LVDMs) have achieved state-of-the-art generative quality for image and video generation; however, they remain brittle under noisy conditioning, where small perturbations in text or multimodal embeddings can cascade over timesteps and cause semantic drift. Existing corruption strategies from image diffusion (Gaussian, Uniform) f
Peiqi Wang, ShengYun Peng, Xuewen Zhang, Hanchao Yu
This work investigates the optimal allocation of inference compute across three key scaling factors in video vision language models: language model size, frame count, and the number of visual tokens per frame. While prior works typically focuses on optimizing model efficiency or improving performance without considering resource constraints, we instead ident
Ashok R Samrat, Swati Singh, Shashi Ranjan Kumar
This paper addresses the time-constrained interception of targets at a predetermined time with bounded field-of-view capability of the seeker-equipped interceptors. We propose guidance laws using the effective lead angle and velocity lead angles of the interceptor to achieve a successful interception of the target. The former scheme extends the existing two-
James L. Martin Robinson, Neshat Moslehi, Nikolaos Dramountanis, Lennart van den Hoven
In biology, ligand mediated transitions (LMT), where the binding of a molecular ligand onto the binding site of a receptor molecule leads to a well-defined change in the conformation of the receptor, are often referred to as 'the second secret of life'. Sharp, cooperative transitions arise in many biological cases, while examples of synthetic cooperative sys
Anzori Sh. Georgadze
This study evaluates the potential of muon tomography as a non-invasive technique for detecting concealed illicit drugs in cargo, based on detailed simulations performed using the GEANT4 toolkit. A combined analysis of muon scattering and absorption data was employed to enhance material discrimination, with a focus on realistic smuggling scenarios involving
The Theory of the Unique Latent Pattern: A Formal Epistemic Framework for Structural Singularity in Complex Systems
cs.AIMohamed Aly Bouke
This paper introduces the Theory of the Unique Latent Pattern (ULP), a formal epistemic framework that redefines the origin of apparent complexity in dynamic systems. Rather than attributing unpredictability to intrinsic randomness or emergent nonlinearity, ULP asserts that every analyzable system is governed by a structurally unique, deterministic generativ
Fractal Attractors in Random Nonlinear Iterated Function Systems: Existence, Stability, and Dimensional Properties
math.DSMohamed Aly Bouke
This study develops a comprehensive theoretical and computational framework for Random Nonlinear Iterated Function Systems (RNIFS), a generalization of classical IFS models that incorporates both nonlinearity and stochasticity. We establish mathematical guarantees for the existence and stability of invariant fractal attractors by leveraging contractivity con
Yichen Luo, Jia Wang, Dapeng Lan, Yu Liu
Partial Differential Equations (PDEs) are fundamental for modeling physical systems, yet solving them in a generic and efficient manner using machine learning-based approaches remains challenging due to limited multi-input and multi-scale generalization capabilities, as well as high computational costs. This paper proposes the Multi-input and Multi-scale Eff
Riccardo Cristoferi, Lorenza D'Elia
We provide a first-order homogenization result for quadratic functionals. In particular, we identify the scaling of the energy and the explicit form of the limiting functional in terms of the first-order correctors. The main novelty of the paper is the use of the dual correspondence between quadratic functionals and PDEs, combined with a refinement of the cl
Signal, Image, or Symbolic: Exploring the Best Input Representation for Electrocardiogram-Language Models Through a Unified Framework
cs.AIWilliam Han, Chaojing Duan, Zhepeng Cen, Yihang Yao
Recent advances have increasingly applied large language models (LLMs) to electrocardiogram (ECG) interpretation, giving rise to Electrocardiogram-Language Models (ELMs). Conditioned on an ECG and a textual query, an ELM autoregressively generates a free-form textual response. Unlike traditional classification-based systems, ELMs emulate expert cardiac elect
Muhammed Golec, Yaser Khamayseh, Suhib Bani Melhem, Abdulmalik Alwarafy
Sixth Generation (6G) wireless networks, which are expected to be deployed in the 2030s, have already created great excitement in academia and the private sector with their extremely high communication speed and low latency rates. However, despite the ultra-low latency, high throughput, and AI-assisted orchestration capabilities they promise, they are vulner
Sagar Sapkota, Mohammad Saqib Hasan, Mubarak Shah, Santu Karmaker
Multi-party Conversational Agents (MPCAs) are systems designed to engage in dialogue with more than two participants simultaneously. Unlike traditional two-party agents, designing MPCAs faces additional challenges due to the need to interpret both utterance semantics and social dynamics. This survey explores recent progress in MPCAs by addressing three key q
M. Pouyez, G. Nicotera, M. Galbiati, T. Grismayer
Electromagnetic showers from high-energy electron beams interacting with a target are a promising path to creating pair plasmas in the laboratory. Here, we solve analytically the kinetic equations describing this process. Two regimes are defined by the ratio of the target thickness $L$ to the shower length $L_{\rm{sh}}$, which depends on the electron energy
Oleg Karpenkov, Brigitte Servatius, Herman Servatius
In this note we provide a topological definition of Maxwell-Cremona liftings for non-planar frameworks of surfaces (both oriented and non-oriented). In the non-oriented case we give an estimate on the dimension of self-stresses, when the frameworks will posses a non-trivial lifting.
Shogo Chiwaki, Ryutaroh Matsumoto
For a quantum secret sharing scheme built from a general quantum stabilizer code, no measurement-free circuit has been known for reconstructing its quantum secrets, except particular classes, such as one proposed by Cleve, Gottesman and Lo. We propose a measurement-free reconstruction circuit of quantum secrets in quantum secret sharing based on stabilizer c
Josh Alman, Shivam Nadimpalli, Shyamal Patel, Rocco A. Servedio
We give two results on PAC learning DNF formulas using membership queries in the challenging "distribution-free" learning framework, where learning algorithms must succeed for an arbitrary and unknown distribution over $\{0,1\}^n$. (1) We first give a quasi-polynomial time "list-decoding" algorithm for learning a single term of an unknown DNF formula. More p
Maicon R. Correa, Abimael F. D. Loula
Unconditionally stable finite element methods for Darcy flow are derived by adding least-squares residual forms of the governing equations to the classical mixed formulations. The proposed methods are free of mesh dependent stabilization parameters and allow the use of the classical continuous Lagrangian finite element spaces of any order for the velocity an
Navojit Dhali Pallab
Pancreatic $\beta-$cells regulate insulin secretion through complex oscillations, which are vital for glucose control and diabetes research. In this paper, an existing mathematical model of $\beta-$cell dynamics is analyzed using a three-time-scale framework to study interactions among fast, intermediate, and slow variables. Through Geometric Singular Pertur
Distributed Incremental SAT Solving with Mallob: Report and Case Study with Hierarchical Planning
cs.DCDominik Schreiber
This report describes an extension of the distributed job scheduling and SAT solving platform Mallob by incremental SAT solving, embedded in a case study on SAT-based hierarchical planning. We introduce a low-latency interface for incremental jobs and specifically for IPASIR-style incremental SAT solving to Mallob. This also allows to process many independen
Paul de Font-Reaulx
Value learning is a crucial aspect of safe and ethical AI. This is primarily pursued by methods inferring human values from behaviour. However, humans care about much more than we are able to demonstrate through our actions. Consequently, an AI must predict the rest of our seemingly complex values from a limited sample. I call this the value generalization p
Samiha Tariq
This paper examines the impact of cognitive biases on financial decision-making through a static Bayesian game framework. While traditional economic theory assumes fully rational investors, real-world choices are often shaped by loss aversion, overconfidence, and herd behavior. Integrating psychological insights with economic game theory, the model studies s
Vignav Ramesh, Morteza Mardani
The iterative and stochastic nature of diffusion models enables test-time scaling, whereby spending additional compute during denoising generates higher-fidelity samples. Increasing the number of denoising steps is the primary scaling axis, but this yields quickly diminishing returns. Instead optimizing the noise trajectory--the sequence of injected noise ve
Johnatan Costa, Ernani Ribeiro, Márcio Santos
In this article, we study quasi-Einstein manifolds with constant scalar curvature. We provide a classification of compact and noncompact (possibly with boundary) $T$-flat quasi-Einstein manifolds with constant scalar curvature, where the $T$-tensor is directly related to the Cotton and Weyl tensors. Moreover, we construct new explicit examples of noncompact
Thomas A. Henzinger, Kaushik Mallik, Pouya Sadeghi, Đorđe Žikelić
We present the first supermartingale certificate for quantitative $\omega$-regular properties of discrete-time infinite-state stochastic systems. Our certificate is defined on the product of the stochastic system and a limit-deterministic B\"uchi automaton that specifies the property of interest; hence we call it a limit-deterministic B\"uchi supermartingale
Arman Zarei, Samyadeep Basu, Keivan Rezaei, Zihao Lin
Understanding how knowledge is distributed across the layers of generative models is crucial for improving interpretability, controllability, and adaptation. While prior work has explored knowledge localization in UNet-based architectures, Diffusion Transformer (DiT)-based models remain underexplored in this context. In this paper, we propose a model- and kn