November 2024 arXiv papers — page 43
Showing 4,201–4,300 of 19,800 papers
Haoyuan Gui, Xiaoyu Zhang, Chong Zhang, Zitong Su
As Convolutional Neural Networks (CNNs) gain prominence in deep learning, algorithms like Winograd Convolution have been introduced to enhance computational efficiency. However, existing implementations often face challenges such as high transformation overhead, suboptimal computation efficiency, and reduced parallel performance in some layers. We propose a
Ishan Panpaliya
The ascending chain condition on principal ideals (ACCP) is almost always complementary to atomicity within integral domains: in fact, Cohn initially stated that these two conditions were equivalent. This assertion has been shown to be false, however most counterexamples require technical algebraic constructions. In 2017, Gotti conjectured that for every $q$
Characterizing Stellar and Gas Properties in NGC 628: Spatial Distributions, Radial Gradients, and Resolved Scaling Relations
astro-ph.GAPeng Wei, Hu Zou, Jing Wang, Xu Kong
Building on our previous research of multi-wavelength data from UV to IR, we employ spectroscopic observations of ionized gas, as well as neutral hydrogen gas obtained from the Five-hundred Meter Aperture Spherical Telescope (FAST), to explore the intrinsic processes of star formation and chemical enrichment within NGC 628. Our analysis focuses on several ke
Niranka Banerjee, Christian Engels, Duc A. Hoang
Reconfiguration problems involve determining whether two given configurations can be transformed into each other under specific rules. The Token Sliding problem asks whether, given two different set of tokens on vertices of a graph $G$, we can transform one into the other by sliding tokens step-by-step along edges of $G$ such that each resulting set of token
Xiangyu Zhu, Chang Yu, Jiankuo Zhao, Zhaoxiang Zhang
David Marr's seminal theory of vision proposes that the human visual system operates through a sequence of three stages, known as the 2D sketch, the 2.5D sketch, and the 3D model. In recent years, Deep Neural Networks (DNN) have been widely thought to have reached a level comparable to human vision. However, the mechanisms by which DNNs accomplish this and w
QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion
cs.SDYoungjun Sim, Jinsung Yoon, Wooyeol Jeong, Young-Joo Suh
Zero-shot voice conversion is a technique that alters the speaker identity of an input speech to match a target speaker using only a single reference utterance, without requiring additional training. Recent approaches extensively utilize self-supervised learning features with K-means quantization to extract high-quality content representations while removing
Wei Lin, Qingyu Song, Hong Xu
Tuning step sizes is crucial for the stability and efficiency of optimization algorithms. While adaptive coordinate-wise step sizes have been shown to outperform scalar step size in first-order methods, their use in second-order methods is still under-explored and more challenging. Current approaches, including hypergradient descent and cutting plane methods
Hsin-Ku Chen, Jheng-Jie Chen, Jungkai A. Chen
We prove that each divisorial contraction to a curve between terminal threefolds is a weighted blow-up under a suitable embedding. Moreover, we give a classification of the weighted blow-ups assuming that the curve is smooth.
Dušica Knežević, Miloš Savić, Miloš Radovanović
The notion of local intrinsic dimensionality (LID) has important theoretical implications and practical applications in the fields of data mining and machine learning. Recent research efforts indicate that LID measures defined for graphs can improve graph representational learning methods based on random walks. In this paper, we discuss how NC-LID, a LID mea
Shijie Pan, Aoran Cheng, Yiqi Sun, Kai Kang
Drone swarms coupled with data intelligence can be the future of wildfire fighting. However, drone swarm firefighting faces enormous challenges, such as the highly complex environmental conditions in wildfire scenes, the highly dynamic nature of wildfire spread, and the significant computational complexity of drone swarm operations. We develop a predict-then
Yifang Hao, Shuchao Li
Given a set of graphs $\mathcal{H}$, we say that a graph $G$ is \textit{$\mathcal{H}$-free} if it does not contain any member of $\mathcal{H}$ as a subgraph. Let $\text{ex}(n,\mathcal{H})$ (resp. $\text{ex}_{sp}(n,\mathcal{H})$) denote the maximum size (resp. spectral radius) of an $n$-vertex $\mathcal{H}$-free graph. Denote by $\text{Ex}(n, \mathcal{H})$ th
Pengfei He
Instruction tuning has become an important step for finetuning pretrained language models to better follow human instructions and generalize on various tasks. Nowadays, pretrained language models become increasingly larger, and full parameter finetuning is overwhelmingly costly. Therefore, Parameter Efficient Finetuning (PEFT) has arisen as a cost-effective
Zhaobin Mo, Qingyuan Liu, Baohua Yan, Longxiang Zhang
Spatiotemporal prediction over graphs (STPG) is crucial for transportation systems. In existing STPG models, an adjacency matrix is an important component that captures the relations among nodes over graphs. However, most studies calculate the adjacency matrix by directly memorizing the data, such as distance- and correlation-based matrices. These adjacency
Andrea Di Lorenzo, Giovanni Inchiostro
We compactify moduli stacks of maps from curves to a large class of quotient stacks $[W/G]$ admitting projective good moduli spaces. Our construction extends quasimap-theoretic compactifications and relies on a new birational operation for algebraic stacks, which we call an extended weighted blow-up. As applications, we construct compactifications of certain
Xiang minjie
Research on Speech Emotion Recognition (SER) often faces challenges such as the lack of large-scale public datasets and limited generalization capability when dealing with data from different distributions. To solve this problem, this paper proposes a cross-corpus speech emotion recognition method based on supervised contrast learning. The method employs a t
Keyu Chen, Wei He, Yixin He, Yuxiang Huang
Recently Shekhar Suman [arXiv: 2407.07121v6 [math.GM] 3 Aug 2024] made an attempt to prove the irrationality of $\zeta(5)$. But unfortunately the proof is not correct. In this note, we discuss the fallacy in the proof.
Zhetao Jia, Hector Rubio, Lilian Neim, Jagang Park
The development of deep neural networks is witnessing fast growth in network size, which requires novel hardware computing platforms with large bandwidth and low energy consumption. Optical computing has been a potential candidate for next-generation computing systems. Specifically, wavelength-division multiplexing (WDM) has been widely adopted in optical ne
Tian Bowen, Lai Songning, Wu Jiemin, Shuai Zhihao
Pretrained models have revolutionized deep learning by enabling significant performance improvements across a wide range of tasks, leveraging large-scale, pre-learned knowledge representations. However, deploying these models in real-world multi-task learning (MTL) scenarios poses substantial challenges, primarily due to high computational costs and ineffici
Feifei Shao, Ping Liu, Zhao Wang, Yawei Luo
Point cloud processing (PCP) encompasses tasks like reconstruction, denoising, registration, and segmentation, each often requiring specialized models to address unique task characteristics. While in-context learning (ICL) has shown promise across tasks by using a single model with task-specific demonstration prompts, its application to PCP reveals significa
QUBO Refinement: Achieving Superior Precision through Iterative Quantum Formulation with Limited Qubits
quant-phHyunju Lee, Kyungtaek Jun
In the era of quantum computing, the emergence of quantum computers and subsequent advancements have led to the development of various quantum algorithms capable of solving linear equations and eigenvalues, surpassing the pace of classical computers. Notably, the hybrid solver provided by the D-wave system can leverage up to two million variables. By exploit
Kousik Giri, Barry Mant, Franco A. Gianturco, Roland Wester
The molecular anion C$_2^-$ has been of interest in the last few years as a candidate for laser cooling due to its electronic structure and favourable branching ratios to the ground electronic and vibrational state. Molecular hydrogen has been used by the Wester group in Innsbruck as a buffer gas to cool the molecule's internal ro-vibrational motion. In the
Can LLMs faithfully generate their layperson-understandable 'self'?: A Case Study in High-Stakes Domains
cs.HCArion Das, Asutosh Mishra, Amitesh Patel, Soumilya De
Large Language Models (LLMs) have significantly impacted nearly every domain of human knowledge. However, the explainability of these models esp. to laypersons, which are crucial for instilling trust, have been examined through various skeptical lenses. In this paper, we introduce a novel notion of LLM explainability to laypersons, termed $\textit{ReQuesting
An Optically Led Search for Kilonovae to z$\sim$0.3 with the Kilonova and Transients Program (KNTraP)
astro-ph.HENatasha Van Bemmel, Jielai Zhang, Jeff Cooke, Armin Rest
Compact binary mergers detectable in gravitational waves can be accompanied by a kilonova, an electromagnetic transient powered by radioactive decay of newly synthesised r-process elements. A few kilonova candidates have been observed during short gamma-ray burst follow-up, and one found associated with a gravitational wave detection, GW170817. However, robu
A Physics-Inspired Deep Learning Framework with Polar Coordinate Attention for Ptychographic Imaging
physics.opticsHan Yue, Jun Cheng, Yu-Xuan Ren, Chien-Chun Chen
Ptychographic imaging confronts inherent challenges in applying deep learning for phase retrieval from diffraction patterns. Conventional neural architectures, both convolutional neural networks and Transformer-based methods, are optimized for natural images with Euclidean spatial neighborhood-based inductive biases that exhibit geometric mismatch with the c
Yu-Kun Zhang, Li-Juan Li, Xue-Ke Song, Liu Ye
An object moving with the acceleration will change the temperature of environment around it, because of the presence of the Unruh thermal effect. In this work, we investigate the impact of Unruh thermal noise on the quantum-memory-assisted {entropic} uncertainty and quantum correlation regarding a pair of Unruh-Dewitt detectors. Specifically, we examine how
Multi-Robot Reliable Navigation in Uncertain Topological Environments with Graph Attention Networks
cs.ROZhuoyuan Yu, Hongliang Guo, Albertus Hendrawan Adiwahono, Jianle Chan
This paper studies the multi-robot reliable navigation problem in uncertain topological networks, which aims at maximizing the robot team's on-time arrival probabilities in the face of road network uncertainties. The uncertainty in these networks stems from the unknown edge traversability, which is only revealed to the robot upon its arrival at the edge's st
Mohammad Hassan Heydari, Arshia Hemmat, Erfan Naman, Afsaneh Fatemi
Retrieval Augmented Generation (RAG) has emerged as a widely adopted approach to mitigate the limitations of large language models (LLMs) in answering domain-specific questions. Previous research has predominantly focused on improving the accuracy and quality of retrieved data chunks to enhance the overall performance of the generation pipeline. However, des
Xinpeng Liu, Hiroaki Santo, Yosuke Toda, Fumio Okura
Accurate estimation of plant skeletal structure (e.g., branching structure) from images is essential for smart agriculture and plant science. Unlike human skeletons with fixed topology, plant skeleton estimation presents a unique challenge, i.e., estimating arbitrary tree graphs from images. While recent graph generation methods successfully infer thin struc
Hyperspectral Image Cross-Domain Object Detection Method based on Spectral-Spatial Feature Alignment
cs.CVHongqi Zhang, He Sun, Hongmin Gao, Feng Han
With consecutive bands in a wide range of wavelengths, hyperspectral images (HSI) have provided a unique tool for object detection task. However, existing HSI object detection methods have not been fully utilized in real applications, which is mainly resulted by the difference of spatial and spectral resolution between the unlabeled target domain and a label
Real-time volumetric free-hand ultrasound imaging for large-sized organs: A study of imaging the whole spine
eess.IVCaozhe Li, Enxiang Shen, Haoyang Wang, Yuxin Wang
Three-dimensional (3D) ultrasound imaging can overcome the limitations of conventional two dimensional (2D) ultrasound imaging in structural observation and measurement. However, conducting volumetric ultrasound imaging for large-sized organs still faces difficulties including long acquisition time, inevitable patient movement, and 3D feature recognition. In
Mahmoud M. Kishky, Hesham M. Eraqi, Khaled F. Elsayed
Autonomous driving involves complex tasks such as data fusion, object and lane detection, behavior prediction, and path planning. As opposed to the modular approach which dedicates individual subsystems to tackle each of those tasks, the end-to-end approach treats the problem as a single learnable task using deep neural networks, reducing system complexity a
A Novel Flow-induced Motion Energy Harvesting with Coupled Mechanism of Time-varying Stiffness and Passive Turbulence Control
physics.flu-dynYongxi Wu, Hao Wu
The ocean contains a substantial amount of energy, and the efficient harvesting of this energy holds significant importance. Drawing inspiration from the biomimicry of octopus tentacles, this study introduces a synergistic mechanism designed to optimize energy harvesting through flowinduced motions, integrating boundary layer modulation via passive turbulenc
Three Cars Approaching within 100m! Enhancing Distant Geometry by Tri-Axis Voxel Scanning for Camera-based Semantic Scene Completion
cs.CVJongseong Bae, Junwoo Ha, Ha Young Kim
Camera-based Semantic Scene Completion (SSC) is gaining attentions in the 3D perception field. However, properties such as perspective and occlusion lead to the underestimation of the geometry in distant regions, posing a critical issue for safety-focused autonomous driving systems. To tackle this, we propose ScanSSC, a novel camera-based SSC model composed
Mohamed Benkedadra, Dany Rimez, Tiffanie Godelaine, Natarajan Chidambaram
Computer vision tasks such as object detection and segmentation rely on the availability of extensive, accurately annotated datasets. In this work, We present CIA, a modular pipeline, for (1) generating synthetic images for dataset augmentation using Stable Diffusion, (2) filtering out low quality samples using defined quality metrics, (3) forcing the existe
Jiahui Liu, Zhenkun Cai, Zhiyong Chen, Minjie Wang
Attention Graph Neural Networks (AT-GNNs), such as GAT and Graph Transformer, have demonstrated superior performance compared to other GNNs. However, existing GNN systems struggle to efficiently train AT-GNNs on GPUs due to their intricate computation patterns. The execution of AT-GNN operations without kernel fusion results in heavy data movement and signif
Vu-Anh Le, Mehmet Dik
We investigate the stability of persistence diagrams \( D \) under non-uniform scaling transformations \( S \) in \( \mathbb{R}^n \). Given a finite metric space \( X \subset \mathbb{R}^n \) with Euclidean distance \( d_X \), and scaling factors \( s_1, s_2, \ldots, s_n > 0 \) applied to each coordinate, we derive explicit bounds on the bottleneck distance \
Kwonjin Park, Jaeyong Cho, Soobeom Lee, Jaehun Cho
Vanadium oxide (VOx) is a material of significant interest due to its metal-insulator transition (MIT) properties as well as its diverse stable antiferromagnetism depending on the valence states of V and O with distinct MIT transitions and N\'eel temperatures. Although several studies reported the ferromagnetism in the VOx, it was mostly associated with impu
Ourania Koutzampasopoulou Xanthidou, Nadine Aburumman, Hanêne Ben-Abdallah
The application of Virtual Reality Environments (VRE) has been gaining momentum as a relatively new tool to assist with mitigating various difficulties including abstractness of concepts, lack of user engagement, perception of disconnection from other users. A VRE may offer both synchronous and asynchronous experiences, in addition to an immersive environmen
Wey Yeh Choong, Yangyang Guo, Mohan Kankanhalli
Vision Large Language Models (VLLMs) are widely acknowledged to be prone to hallucinations. Existing research addressing this problem has primarily been confined to image inputs, with limited exploration of video-based hallucinations. Furthermore, current evaluation methods fail to capture nuanced errors in generated responses, which are often exacerbated by
Med-PerSAM: One-Shot Visual Prompt Tuning for Personalized Segment Anything Model in Medical Domain
cs.CVHangyul Yoon, Doohyuk Jang, Jungeun Kim, Eunho Yang
Leveraging pre-trained models with tailored prompts for in-context learning has proven highly effective in NLP tasks. Building on this success, recent studies have applied a similar approach to the Segment Anything Model (SAM) within a ``one-shot" framework, where only a single reference image and its label are employed. However, these methods face limitatio
From Collapse to Stability: A Knowledge-Driven Ensemble Framework for Scaling Up Click-Through Rate Prediction Models
cs.IRHonghao Li, Lei Sang, Yi Zhang, Guangming Cui
Click-through rate (CTR) prediction plays a crucial role in modern recommender systems. While many existing methods utilize ensemble networks to improve CTR model performance, they typically restrict the ensemble to only two or three sub-networks. Whether increasing the number of sub-networks consistently enhances CTR model performance to align with scaling
DP-CDA: An Algorithm for Enhanced Privacy Preservation in Dataset Synthesis Through Randomized Mixing
stat.MLUtsab Saha, Tanvir Muntakim Tonoy, Hafiz Imtiaz
In recent years, the growth of data across various sectors, including healthcare, security, finance, and education, has created significant opportunities for analysis and informed decision-making. However, these datasets often contain sensitive and personal information, which raises serious privacy concerns. It has been shown in multiple works that a person'
Rui Zuo, Simon Khan, Zifan Wang, Garrett Ethan Katz
Reinforcement learning (RL) has demonstrated remarkable success in solving complex decision-making problems, yet its adoption in critical domains is hindered by the lack of interpretability in its decision-making processes. Existing explainable AI (xAI) approaches often fail to provide meaningful explanations for RL agents, particularly because they overlook
Diederik Aerts, Massimiliano Sassoli de Bianchi, Sandro Sozzo
An overview of the conceptuality interpretation of quantum mechanics is presented, along with an explanation of how it sheds light on key quantum and relativistic phenomena. In particular, we show how the interpretation clarifies Heisenberg's uncertainty principle, wave function-based and entanglement-based nonlocality, interference effects resulting from th
Xi Zhang, Xiaolin Wu
It is customary to deploy uniform scalar quantization in the end-to-end optimized Neural image compression methods, instead of more powerful vector quantization, due to the high complexity of the latter. Lattice vector quantization (LVQ), on the other hand, presents a compelling alternative, which can exploit inter-feature dependencies more effectively while
Comparative Analysis of Machine Learning Models for Short-Term Distribution System Load Forecasting
eess.SYElias Raffoul, Mingjian Tuo, Cunzhi Zhao, Tianxia Zhao
Accurate electrical load forecasting is crucial for optimizing power system operations, planning, and management. As power systems become increasingly complex, traditional forecasting methods may fail to capture the intricate patterns and dependencies within load data. Machine learning (ML) techniques have emerged as powerful alternatives, offering superior
Yuji Cao, Yue Chen, Yan Xu
The stochastic nature of renewable energy and load demand requires efficient and accurate solutions for probabilistic optimal power flow (OPF). Quantum neural networks (QNNs), which combine quantum computing and machine learning, offer computational advantages in approximating OPF by effectively handling high-dimensional data. However, adversaries with acces
Raquib Bin Yousuf, Nicholas Defelice, Mandar Sharma, Shengzhe Xu
Building on their demonstrated ability to perform a variety of tasks, we investigate the application of large language models (LLMs) to enhance in-depth analytical reasoning within the context of intelligence analysis. Intelligence analysts typically work with massive dossiers to draw connections between seemingly unrelated entities, and uncover adversaries'
Monika, Kirandeep Kaur, Varinder Singh, Shishram Rebari
We study a quantum harmonic Otto engine under a hot squeezed thermal reservoir with asymmetry between the two adiabatic branches introduced by considering different speeds of the driving protocols. In the first configuration, the driving protocol for the expansion stroke is sudden-switch in nature and compression stroke is driven adiabatically, while the sec
Chen Tan, M. Le Delliou, Ke Wang
According to the Schr\"odinger-Poisson equations, fuzzy dark matter (FDM) can form a stable equilibrium configuration, the so-called FDM soliton. In principle, given the FDM particle mass, the profile of the FDM soliton is fixed. In practice, however, there is a great diversity of structures in the Universe. Possible causes of such diversity can lie in such
Ira M. Gessel
Answering a question of Donald Knuth, we find the bivariate exponential generating function for "up-up-or-down-down'' permutations of odd length according to their last entry. An up-up-or-down-down permutation is a permutation $a_1a_2\cdots a_n$ satisfying $a_{2i-1}<a_{2i}$ if and only if $a_{2i}<a_{2i+1}$ for $1\le i <n/2$. Equivalently, an up-up-or-down-do
Hao Chang, Hoang Triet Vo, Alva Kosasih, Branka Vucetic
End-to-end (E2E) learning has recently been proposed to jointly design the modulator and symbol detector by using deep neural networks (DNNs). However, existing schemes lack sufficient capability to cancel multi-user interference (MUI) in uplink multi-user multiple-input multiple-output (MU-MIMO) systems. In this paper, we propose a graph neural network (GNN
Vasudev Gohil, Matthew DeLorenzo, Veera Vishwa Achuta Sai Venkat Nallam, Joey See
The rapid advancement of large language models (LLMs) has enabled the ability to effectively analyze and generate code nearly instantaneously, resulting in their widespread adoption in software development. Following this advancement, researchers and companies have begun integrating LLMs across the hardware design and verification process. However, these hig
Jiin Im, Yongho Son, Je Hyeong Hong
While the mainstream research in anomaly detection has mainly followed the one-class classification, practical industrial environments often incur noisy training data due to annotation errors or lack of labels for new or refurbished products. To address these issues, we propose a novel learning-based approach for fully unsupervised anomaly detection with unl
Tangli Ge
In this paper, we establish the following family version of Habegger's bounded height theorem on abelian varieties: a locally closed subvariety of an abelian scheme with Gao's $t^{\mathrm{th}}$ degeneracy locus removed, intersected with all flat group subschemes of relative dimension at most $t$, gives a set of bounded total height. Our main tools include th
Nathaniel Hanson, Sarvesh Prajapati, James Tukpah, Yash Mewada
With the rapid increase in wildfires in the past decade, it has become necessary to detect and predict these disasters to mitigate losses to ecosystems and human lives. In this paper, we present a novel solution -- Hyper-Drive3D -- consisting of snapshot hyperspectral imaging and LiDAR, mounted on an Unmanned Ground Vehicle (UGV) that identifies areas inside
Xingyu Liu, Gu Wang, Ruida Zhang, Chenyangguang Zhang
Unseen object pose estimation methods often rely on CAD models or multiple reference views, making the onboarding stage costly. To simplify reference acquisition, we aim to estimate the unseen object's pose through a single unposed RGB-D reference image. While previous works leverage reference images as pose anchors to limit the range of relative pose, our s
Jatin Nainani, Sankaran Vaidyanathan, AJ Yeung, Kartik Gupta
Mechanistic interpretability aims to understand the inner workings of large neural networks by identifying circuits, or minimal subgraphs within the model that implement algorithms responsible for performing specific tasks. These circuits are typically discovered and analyzed using a narrowly defined prompt format. However, given the abilities of large langu
You only thermoelastically deform once: Point Absorber Detection in LIGO Test Masses with YOLO
astro-ph.IMSimon R. Goode, Mitchell Schiworski, Daniel Brown, Eric Thrane
Current and future gravitational-wave observatories rely on large-scale, precision interferometers to detect the gravitational-wave signals. However, microscopic imperfections on the test masses, known as point absorbers, cause problematic heating of the optic via absorption of the high-power laser beam, which results in diminished sensitivity, lock loss, or
Non-commutative Stein's Method: Applications to Free Probability and Sums of Non-commutative Variables
math.PRMario Díaz, Arturo Jaramillo
We present a straightforward formulation of Stein's method for the semicircular distribution, specifically designed for the analysis of non-commutative random variables. Our approach employs a non-commutative version of Stein's heuristic, interpolating between the target and approximating distributions via the free Ornstein-Uhlenbeck semigroup. A key applica
Rikhav Shah, Isabel Detherage
This paper asks if the following iterative procedure approximately orthogonalizes a set of $n$ linearly independent unit vectors while preserving their span: in each iteration, access a random pair of vectors and replace one with the component perpendicular to the other, renormalized to be a unit vector. We provide a positive answer: any given set of startin
Christoph Treude, Christopher M. Poskitt
As software development increasingly adopts automation, bot-driven development (BotDD) represents a transformative shift where bots assume proactive roles in coding, testing, and project management. In bot-driven development, bots go beyond support tasks, actively driving development workflows by making autonomous decisions, performing independent assessment
Peiheng Zhou, Ming Hu, Xingrun Quan, Yawen Peng
Although Deep Learning (DL) methods becoming increasingly popular in vulnerability detection, their performance is seriously limited by insufficient training data. This is mainly because few existing software organizations can maintain a complete set of high-quality samples for DL-based vulnerability detection. Due to the concerns about privacy leakage, most
Diliang Chen, Nozhan Ghoreishi, John LaCourse, Sajay Arthanat
Lifting during manual material handling is a major cause of low-back pain (LBP). As an important risk factor that directly influences the risk of LBP, the Load vertical location (LVL) during lifting needs to be measured and controlled. However, existing solutions for LVL measurement are inefficient, inaccurate, and impractical for real-world workplace enviro
ENCLIP: Ensembling and Clustering-Based Contrastive Language-Image Pretraining for Fashion Multimodal Search with Limited Data and Low-Quality Images
cs.CVPrithviraj Purushottam Naik, Rohit Agarwal
Multimodal search has revolutionized the fashion industry, providing a seamless and intuitive way for users to discover and explore fashion items. Based on their preferences, style, or specific attributes, users can search for products by combining text and image information. Text-to-image searches enable users to find visually similar items or describe prod
Peng Cui, Yiming Yang, Fusheng Jin, Siyuan Tang
In online advertising, once an ad campaign is deployed, the automated bidding system dynamically adjusts the bidding strategy to optimize Cost Per Action (CPA) based on the number of ad conversions. For ads with a long conversion delay, relying solely on the real-time tracked conversion number as a signal for bidding strategy can significantly overestimate t
Tatsuya Yokota
Tensor network diagram (graphical notation) is a useful tool that graphically represents multiplications between multiple tensors using nodes and edges. Using the graphical notation, complex multiplications between tensors can be described simply and intuitively, and it also helps to understand the essence of tensor products. In fact, most of matrix/tensor p
SiGe BiCMOS Circuit Design using only PMOS and HBTs Approach for the Ocean Worlds Exploration
physics.ins-detMd Omar Faruk, Steven Corum, Zakaraya Hamdan, Alex Seaver
Space exploration to have the biosignatures of extraterrestrial life on different planets with oceans in our solar system and beyond requires the design and manufacturing of robust and reliable electronic systems that can be used for sensing, data processing, controlling motor/actuators, and communication while surviving an extreme environment. Commercial of
Fundamental Microscopic Properties as Predictors of Large-Scale Quantities of Interest: Validation through Grain Boundary Energy Trends
cond-mat.mtrl-sciBenjamin A. Jasperson, Ilia Nikiforov, Amit Samanta, Brandon Runnels
Correlations between fundamental microscopic properties computable from first principles, which we term canonical properties, and complex large-scale quantities of interest (QoIs) provide an avenue to predictive materials discovery. We propose that such correlations can be efficiently discovered through simulations utilizing approximate interatomic potential
Oki Gunawan, Chaeyoun Kim, Bonfilio Nainggolan, Minyeul Lee
Electronic trap states are a critical yet unavoidable aspect of semiconductor devices, impacting performance of various electronic devices such as transistors, memory devices, solar cells, and LEDs. The density, energy level, and position of these trap states often enable or constrain device functionality, making their measurement crucial in materials scienc
Biao Wu, Huajun Zhang
Two families $\mathcal{A}$ and $\mathcal{B}$ of sets are called cross-intersecting if each pair of sets $A\in \mathcal{A}$ and $B\in \mathcal{B}$ has nonempty intersection. Let $\cal{A}$ and ${\cal B}$ be two cross-intersecting families of $k$-subsets and $\ell$-subsets of $[n]$. Matsumoto and Tokushige [J. Combin. Theory Ser. A 52 (1989) 90--97] studied the
Andrew Hassell, Qiuye Jia
In this article we consider the defocusing nonlinear Schr\"odinger equation, with time-dependent potential, in space dimensions $n=1, 2$ and $3$, with nonlinearity $|u|^{p-1} u$, $p$ an odd integer, satisfying $p \geq 5$ in dimension $1$, $p \geq 3$ in dimension $2$ and $p=3$ in dimension $3$. We also allow a metric perturbation, assumed to be compactly supp
The Frequency and Mass-Ratio Distribution of Binaries in Clusters -- III: Probabilistic Generative Modelling of Six Young Open Clusters
astro-ph.GAJason Alexander, Michael Albrow
We apply probabilistic generative modelling of colour-magnitude diagrams to six young Galactic open star clusters and determine their mass functions, binary mass-ratio distributions, and the frequencies of binary stars. We find that younger clusters tend to exhibit a higher incidence of binaries than their older counterparts. The mass-ratio distribution is f
Abhishek Yadav, Francesco Caravelli, David Wolpert
Mismatch cost (MMC) is a universally applicable lower bound on the entropy production (EP) of any fixed physical process across a given time interval. In the first part of the paper, we establish results concerning MMC to prove that it scales at least linearly with the total heat flow in the worst case over initial distributions. We also prove that the MMC l
AI-Generated Image Quality Assessment Based on Task-Specific Prompt and Multi-Granularity Similarity
cs.CVJili Xia, Lihuo He, Fei Gao, Kaifan Zhang
Recently, AI-generated images (AIGIs) created by given prompts (initial prompts) have garnered widespread attention. Nevertheless, due to technical nonproficiency, they often suffer from poor perception quality and Text-to-Image misalignment. Therefore, assessing the perception quality and alignment quality of AIGIs is crucial to improving the generative mod
Zain Taufique, Aman Vyas, Antonio Miele, Pasi Liljeberg
Edge inference techniques partition and distribute Deep Neural Network (DNN) inference tasks among multiple edge nodes for low latency inference, without considering the core-level heterogeneity of edge nodes. Further, default DNN inference frameworks also do not fully utilize the resources of heterogeneous edge nodes, resulting in higher inference latency.
Kaizhao Liang, Lizhang Chen, Bo Liu, Qiang Liu
AdamW has been the default optimizer for transformer pretraining. For many years, our community searched for faster and more stable optimizers with only constrained positive outcomes. In this work, we propose a \textbf{one-line modification in Pytorch} to any momentum-based optimizer, which we rename cautious optimizer, e.g. C-AdamW and C-Lion. Our theoretic
Shuyan Cheng, Yishu Wei, Yiliang Zhou, Zihan Xu
Objectives: The vast and complex nature of human genomic sequencing data presents challenges for effective analysis. This review aims to investigate the application of Natural Language Processing (NLP) techniques, particularly Large Language Models (LLMs) and transformer architectures, in deciphering genomic codes, focusing on tokenization, transformer model
Data Processing Efficiency Aware User Association and Resource Allocation in Blockchain Enabled Metaverse over Wireless Communications
eess.SPLiangxin Qian, Jun Zhao
In the rapidly evolving landscape of the Metaverse, enhanced by blockchain technology, the efficient processing of data has emerged as a critical challenge, especially in wireless communication systems. Addressing this need, our paper introduces the innovative concept of data processing efficiency (DPE), aiming to maximize processed bits per unit of resource
Haojie Huang, Hongchen Luo, Wei Zhai, Yang Cao
Intelligent agents accomplish different tasks by utilizing various objects based on their affordance, but how to select appropriate objects according to task context is not well-explored. Current studies treat objects within the affordance category as equivalent, ignoring that object affordances vary in priority with different task contexts, hindering accura
Congliang Chen, Li Shen, Zhiqiang Xu, Wei Liu
Bi-level optimization has achieved considerable success in contemporary machine learning applications, especially for given proper hyperparameters. However, due to the two-level optimization structure, commonly, researchers focus on two types of bi-level optimization methods: approximate implicit differentiation (AID)-based and iterative differentiation (ITD
Yitong Wang, Xudong Xu, Li Ma, Haoran Wang
Automatic 3D content creation has gained increasing attention recently, due to its potential in various applications such as video games, film industry, and AR/VR. Recent advancements in diffusion models and multimodal models have notably improved the quality and efficiency of 3D object generation given a single RGB image. However, 3D objects generated even
Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting
cs.LGZhi-Yi Chin, Pin-Yu Chen, Wei-Chen Chiu, Mario Fritz
Understanding the capabilities of text-to-image (T2I) models in harmful content generation is essential to safety and compliance. However, human red-teaming is costly and inconsistent, driving the need for automatic tools that simulate realistic misuse attempts. Existing methods either require white-box access, fail to generalize across defenses, or produce
Donggeun Ko, Dongjun Lee, Namjun Park, Wonkyeong Shim
Neural networks struggle with image classification when biases are learned and misleads correlations, affecting their generalization and performance. Previous methods require attribute labels (e.g. background, color) or utilizes Generative Adversarial Networks (GANs) to mitigate biases. We introduce DiffuBias, a novel pipeline for text-to-image generation th
Wenlin Qiu, Tao Guo, Yiqun Li, Xu Guo
We consider the variable-exponent Abel kernel and demonstrate its multiscale nature in modeling crossover dynamics from the initial quasi-exponential behavior to long-term power-law behavior. Then we apply this to an integro-differential equation modeling, e.g. mechanical vibration of viscoelastic materials with changing material properties. We apply the Cra
SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text
cs.CLReshmi Ghosh, Tianyi Yao, Lizzy Chen, Sadid Hasan
Large Language Model (LLM) integrations into applications like Microsoft365 suite and Google Workspace for creating/processing documents, emails, presentations, etc. has led to considerable enhancements in productivity and time savings. But as these integrations become more more complex, it is paramount to ensure that the quality of output from the LLM-integ
Biao Zhang, Jing Ren, Peter Wonka
Neural representations of 3D data have been widely adopted across various applications, particularly in recent work leveraging coordinate-based networks to model scalar or vector fields. However, these approaches face inherent challenges, such as handling thin structures and non-watertight geometries, which limit their flexibility and accuracy. In contrast,
The brain versus AI: World-model-based versatile circuit computation underlying diverse functions in the neocortex and cerebellum
q-bio.NCShogo Ohmae, Keiko Ohmae
AI's significant recent advances using general-purpose circuit computations offer a potential window into how the neocortex and cerebellum of the brain are able to achieve a diverse range of functions across sensory, cognitive, and motor domains, despite their uniform circuit structures. However, comparing the brain and AI is challenging unless clear similar
Wangze Xu, Yifan Zhan, Zhihang Zhong, Xiao Sun
The emergence of neural rendering has significantly advanced the rendering quality of 3D human avatars, with the recently popular 3DGS technique enabling real-time performance. However, SMPL-driven 3DGS human avatars still struggle to capture fine appearance details due to the complex mapping from pose to appearance during fitting. In this paper, we propose
Youngjae Cho, Gwangyeol Kim, Sirojbek Safarov, Seongdeok Bang
Detecting anomalies in industrial settings is challenging due to the scarcity of labeled anomalous data. Generative models can mitigate this issue by synthesizing realistic defect samples, but existing approaches often fail to model the crucial interplay between defects and their background. This oversight leads to unrealistic anomalies, especially in scenar
Sepehr Moalemi, James Richard Forbes
This paper presents a discrete-time passivity-based analysis of the gradient descent method for a class of functions with sector-bounded gradients. Using a loop transformation, it is shown that the gradient descent method can be interpreted as a passive controller in negative feedback with a very strictly passive system. The passivity theorem is then used to
Sw. Banerjee, E. Ben-Haim, F. Bernlochner, E. Bertholet
This paper reports world averages of measurements of $b$-hadron, $c$-hadron, and $\tau$-lepton properties obtained by the Heavy Flavour Averaging Group using results available before October 2023. In rare cases, significant results obtained several months later are also used. For the averaging, common input parameters used in the various analyses are adjuste
Zhu Yu, Bowen Pang, Lizhe Liu, Runmin Zhang
We introduce LOcc, an effective and generalizable framework for open-vocabulary occupancy (OVO) prediction. Previous approaches typically supervise the networks through coarse voxel-to-text correspondences via image features as intermediates or noisy and sparse correspondences from voxel-based model-view projections. To alleviate the inaccurate supervision,
Ovidiu Costin, Rodica Costin, Kriti Sehgal
The H\'enon-Heiles system, initially introduced as a simplified model of galactic dynamics, has become a paradigmatic example in the study of nonlinear systems. Despite its simplicity, it exhibits remarkably rich dynamical behavior, including the interplay between regular and chaotic orbital dynamics, resonances, and stochastic regions in phase space, which
Chang Bi, Kailun Bai, Xing Li, Xuekui Zhang
We introduce HiCat (Hybrid Cell Annotation using Transformative embeddings), a novel semi-supervised pipeline for annotating cell types from single-cell RNA sequencing data. HiCat fuses the strengths of supervised learning for known cell types with unsupervised learning to identify novel types. This hybrid approach incorporates both reference and query genom
The fate of EMRI-IMRI pairs in AGN accretion disks: hydrodynamic and three body simulations
astro-ph.HEPeng Peng, Alessia Franchini, Matteo Bonetti, Alberto Sesana
Extreme-mass-ratio inspirals (EMRIs) and intermediate-mass-ratio inspirals (IMRIs) are important gravitational wave (GW) sources for the Laser Interferometer Space Antenna (LISA). It has been recently suggested that EMRIs and IMRIs can both form in the accretion disk of an active galactic nucleus (AGN). Considering the likely encounter between a sBH and an I
Shayne Waldron
We give a simple presentation of the six quaternionic equiangular lines in $\mathbb{H}^2$ as an orbit of the primitive quaternionic reflection group of order 720 (which is isomorphic to 2.A_6 the double cover of $A_6)$. Other orbits of this group are also seen to give optimal spherical designs (packings) of 10, 15 and 20 lines in $\mathbb{H}^2$, with angles
Yuetong Zhao, Wenguang Zhai
Let N be a large enough natural number, A and B be subsets of {N+1, ... , 2N}. In this paper, we prove that there exists integers a, b with a belongs to A, b belongs to B such that ab=P_k^2 + O(P_k^{1-c}), where 0<c<1/2 and P_k denotes an almost-prime with at most k prime factors, counted with multiplicity.
Generation of circular field harmonics in quasi-polygonal magnet apertures using superconducting canted-cosine-theta coils
physics.acc-phJie Li, Kedong Wang, Kun Zhu
Superconducting magnets with non-circular apertures are important for handling unconventional beam profiles and specialized accelerator applications. This paper presents an analytical framework for designing superconducting accelerator magnets with quasi-polygonal apertures, aimed at generating precise circular field harmonics. In Part 1, we explore the rela
A priori and a posteriori error estimates of a really pressure-robust virtual element method for the incompressible Brinkman problem
math.NAYu Xiong, Yanping Chen
This paper presents both a priori and a posteriori error analyses for a really pressure-robust virtual element method to approximate the incompressible Brinkman problem. We construct a divergence-preserving reconstruction operator using the Raviart-Thomas element for the discretization on the right-hand side. The optimal priori error estimates are carried ou