November 2024 arXiv papers — page 47
Showing 4,601–4,700 of 19,800 papers
Zhong-Yu Li, Yu-Song Hu, Bo-Wen Yin, Ming-Ming Cheng
Vision representation learning, especially self-supervised learning, is pivotal for various vision applications. Ensemble learning has also succeeded in enhancing the performance and robustness of the vision models. However, traditional ensemble strategies are impractical for representation learning, especially self-supervised representation learning that re
A reassessment of LVE method and hemispherical power asymmetry in CMB temperature data from Planck PR4
astro-ph.COSanjeev Sanyal, Sanjeet K. Patel, Pavan K. Aluri, Arman Shafieloo
We undertake a reassessment of one of the large angular scale anomalies observed in cosmic microwave background (CMB) temperature signal referred to as Hemispherical Power Asymmetry (HPA). For the present analysis we use \texttt{SEVEM} cleaned CMB maps from \emph{Planck}'s 2020 final data release (public release 4/PR4). To probe HPA, we employed the local va
Zhonghua Yi, Ge Niu, Lei Wang, Wei Tang
This paper introduces a novel approach, the Bounded-Cache Transformer (BCT), for building large language models with a predefined Key-Value (KV) cache capacity. The BCT addresses the excessive memory consumption issue in traditional KV caches by implementing a bounded-length KV cache, which is particularly suitable for the attention layers in Transformer dec
Kunal Pandey, Rathin Adhikari
A novel scenario is presented within the Type-I seesaw mechanism in which no other beyond Standard Model fields except three heavy right handed neutrinos, have been considered. Light neutrino masses around sub eV scale, could be possible at low seesaw scale around TeV or even below that. At the leading order, 6x6 seesaw mass matrix reproduces three massless
Keunwoo Park, Subin Ahn, Mina Jung, You Jung Cho
Play is a fundamental aspect of developmental growth, yet many parents encounter significant challenges in fulfilling their caregiving roles in this area. As online content increasingly serves as the primary source of parental guidance, this study investigates the difficulties parents face related to play and evaluates the limitations of current online conte
Muhammad Ali Raza, Francisco Tello-Ortiz, M. Zubair, Y. Gómez-Leyton
In this work, we consider a static wormhole in Bopp-Podolsky electrodynamics and convert it into its rotating counterpart by reducing it into Morris-Thorne form. We further study the null geodesics and effective potential along with the shadows for inner and outer unstable orbits for specific choices of parameters. It is found that for some cases smooth shad
Wanting Yang, Zehui Xiong, Song Guo, Shiwen Mao
With the impressive generative capabilities of diffusion models, personalized content synthesis has emerged as the most highly anticipated. However, the large model sizes and iterative nature of inference make it difficult to deploy personalized diffusion models broadly on local devices with varying computational power. To this end, we propose a novel framew
Distinctive Electronic Characteristics and Ultra-high Thermoelectric Power Factor in Be-Fe Intermetallics
cond-mat.mtrl-sciQ. D. Hao, H. Wang, X. R. Chen, Hua Y. Geng
Beryllium (Be) alloys are indispensable in cutting-edge applications due to their unique advantages. However, the scientific understanding about their structure and property is deficient, which greatly restricts their applications within a narrow field. In this work, a systematic investigation on the structure and properties of Be-Fe binary was carried out w
Yu Chen, Rolandos Alexandros Potamias, Evangelos Ververas, Jifei Song
Neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) are popular techniques to reconstruct and render photo-realistic images. However, the pre-requisite of running Structure-from-Motion (SfM) to get camera poses limits their completeness. While previous methods can reconstruct from a few unposed images, they are not applicable when images are unord
Yassine Machta, Omar Ali, Kevin Hakkakian, Ana Vlasceanu
Surgical assessment of liver cancer patients requires identification of the vessel trees from medical images. Specifically, the venous trees - the portal (perfusing) and the hepatic (draining) trees are important for understanding the liver anatomy and disease state, and perform surgery planning. This research aims to improve the 3D segmentation, skeletoniza
Álvaro Navarrete, Víctor Zapatero, Marcos Curty
Recent advancements in quantum key distribution have led to the development of various modulator-free transmitters. Among their advantages, these transmitters offer enhanced security against Trojan-horse attacks. However, practical implementations emit residual pulses that, while not used in the quantum communication, still carry information about Alice's se
Qia Li, Na Zhang, Junyu Feng, Hanwei Yan
In this paper, we consider a class of structured nonsmooth optimization problems over an embedded submanifold of a Euclidean space, where the first part of the objective is the sum of a difference-of-convex (DC) function and a smooth function, while the remaining part is a weakly convex function over a smooth function. This model problem has many important a
A robust time-split linearized explicit/implicit technique for two-dimensional hydrodynamic model: an application to floods in Cameroon far north region
math.NAEric Ngondiep
This paper deals with a time-split explicit/implicit approach for solving a two-dimensional hydrodynamic flow model with appropriate initial and boundary conditions. The time-split technique is employed to upwind the convection term and to treat the friction slope so that the numerical oscillations and stability are well controlled. A suitable time step rest
Timo Eckhardt, David J. Pym
Proof-theoretic semantics, and base-extension semantics in particular, can be seen as a logical realization of inferentialism, in which the meaning of expressions is understood through their use. We present a base-extension semantics for public announcement logic, building on earlier work giving a base-extension semantics for the modal logic $S5$, which in t
Jagadish Pine
In this short note, we provide an alternative proof of a notable theorem by Narasimhan and Ramanan. The theorem states that the moduli space of $S$-equivalence classes of semistable rank $2$ vector bundles over a curve $X$ of genus $2$ with trivial determinant is isomorphic to $\mathbb{P}^3$. Our proof relies on a criterion by Bauer and Szemberg, which chara
Filza Akhlaq, Alina Arshad, Muhammad Yehya Hayati, Jawwad A. Shamsi
Detecting mixed-critical events through computer vision is challenging due to the need for contextual understanding to assess event criticality accurately. Mixed critical events, such as fires of varying severity or traffic incidents, demand adaptable systems that can interpret context to trigger appropriate responses. This paper addresses these challenges b
Chenglong Liu, Jintao Liu, Haorao Wei, Jinze Yang
The corner-based detection paradigm enjoys the potential to produce high-quality boxes. But the development is constrained by three factors: 1) Hard to match corners. Heuristic corner matching algorithms can lead to incorrect boxes, especially when similar-looking objects co-occur. 2) Poor instance context. Two separate corners preserve few instance semantic
Muyao Niu, Yifan Zhan, Qingtian Zhu, Zhuoxiao Li
The creation of 3D human avatars from multi-view videos is a significant yet challenging task in computer vision. However, existing techniques rely on high-quality, sharp images as input, which are often impractical to obtain in real-world scenarios due to variations in human motion speed and intensity. This paper introduces a novel method for directly recon
Zhicheng Zhao, Changfu Zhou, Yu Zhang, Chenglong Li
Remote Sensing Visual Question Answering (RSVQA) has gained significant research interest. However, current RSVQA methods are limited by the imaging mechanisms of optical sensors, particularly under challenging conditions such as cloud-covered and low-light scenarios. Given the all-time and all-weather imaging capabilities of Synthetic Aperture Radar (SAR),
Gradient Norm Regularization Second-Order Algorithms for Solving Nonconvex-Strongly Concave Minimax Problems
math.OCJun-Lin Wang, Zi Xu
In this paper, we study second-order algorithms for solving nonconvex-strongly concave minimax problems, which have attracted much attention in recent years in many fields, especially in machine learning.We propose a gradient norm regularized trust-region (GRTR) algorithm to solve nonconvex-strongly concave minimax problems, where the objective function of t
Umur Togay Yazar, Mucahid Kutlu
Dynamic structure of languages poses significant challenges in applying natural language processing models on historical texts, causing decreased performance in various downstream tasks. Turkish is a prominent example of rapid linguistic transformation due to the language reform in the 20th century. In this paper, we propose two methods for detecting synonym
Simultaneous Measurement of Thermal Conductivity, Heat Capacity, and Interfacial Thermal Conductance by Leveraging Negative Delay-Time Data in Time-Domain Thermoreflectance
cond-mat.mtrl-sciMingzhen Zhang, Tao Chen, Ao Zeng, Jialin Tang
Time-domain thermoreflectance (TDTR) is a widely used technique for characterizing the thermal properties of bulk and thin-film materials. Traditional TDTR analyses typically focus on positive delay time data for fitting, often requiring multiple-frequency measurements to simultaneously determine thermal conductivity and heat capacity. However, this multiple
Suyuan Huang, Chao Zhang, Yuanyuan Wu, Haoxin Zhang
Dense retrieval in most industries employs dual-tower architectures to retrieve query-relevant documents. Due to online deployment requirements, existing real-world dense retrieval systems mainly enhance performance by designing negative sampling strategies, overlooking the advantages of scaling up. Recently, Large Language Models (LLMs) have exhibited super
Yaser Alizadeh, Sandi Klavžar, Javaher Langari
For a positive integer $k\ge 1$, a graph $G$ is $k$-stepwise irregular ($k$-SI graph) if the degrees of every pair of adjacent vertices differ by exactly $k$. Such graphs are necessarily bipartite. Using graph products it is demonstrated that for any $k\ge 1$ and any $d \ge 2$ there exists a $k$-SI graph of diameter $d$. A sharp upper bound for the maximum d
Yi Yan, Dayu Qin, Ercan Engin Kuruoglu
This work introduces the LLM Online Spatial-temporal Reconstruction (LLM-OSR) framework, which integrates Graph Signal Processing (GSP) and Large Language Models (LLMs) for online spatial-temporal signal reconstruction. The LLM-OSR utilizes a GSP-based spatial-temporal signal handler to enhance graph signals and employs LLMs to predict missing values based o
Arvind Murari Vepa, Zukang Yang, Andrew Choi, Jungseock Joo
Deep learning has seen remarkable advancements in machine learning, yet it often demands extensive annotated data. Tasks like 3D semantic segmentation impose a substantial annotation burden, especially in domains like medicine, where expert annotations drive up the cost. Active learning (AL) holds great potential to alleviate this annotation burden in 3D med
Deriving Tsallis entropy from non-extensive Hamiltonian within a statistical mechanics framework
cond-mat.stat-mechParadon Krisut, Sikarin Yoo-Kong
The Tsallis entropy, which possesses non-extensive property, is derived from the first principle employing the non-extensive Hamiltonian or the $q$-deformed Hamiltonian with the canonical ensemble assumption in statistical mechanics. Here, the $q$-algebra and properties of $q$-deformed functions are extensively used throughout the derivation. Consequently, t
Yifan Guo
Thanks to the low cost and power consumption, hybrid analog-digital architectures are considered as a promising energy-efficient solution for massive multiple-input multiple-output (MIMO) systems. The key idea is to connect one RF chain to multiple antennas through low-cost phase shifters. However, due to the non-convex objective function and constraints, we
Chunhui Zhang, Li Liu, Hao Wen, Xi Zhou
Night unmanned aerial vehicle (UAV) tracking is impeded by the challenges of poor illumination, with previous daylight-optimized methods demonstrating suboptimal performance in low-light conditions, limiting the utility of UAV applications. To this end, we propose an efficient mamba-based tracker, leveraging dual enhancement techniques to boost night UAV tra
Chaviva Sirote-Katz, Ofri Palti, Naomi Spiro, Tamás Kálmán
Combinatorial mechanical metamaterials are made of anisotropic, flexible blocks, such that multiple metamaterials may be constructed using a single block type, and the system's response strongly depends on the mutual orientations of the blocks within the lattice. We study a family of possible block types for the square, honeycomb, and cubic lattices. Blocks
Muhammad Suleman Ali Hamdani, Khizer Zakir, Neetu Kushwaha, Syeda Eman Fatima
Brick kilns are a major source of air pollution in Pakistan, with many operating without regulation. A key challenge in Pakistan and across the Indo-Gangetic Plain is the limited air quality monitoring and lack of transparent data on pollution sources. To address this, we present a two-fold AI approach that combines low-resolution Sentinel-2 and high-resolut
Yanchen Zhao, Wenhong Duan, Chuanmin Jia, Shanshe Wang
In the fourth generation Audio Video coding Standard (AVS4), the Inter Prediction Filter (INTERPF) reduces discontinuities between prediction and adjacent reconstructed pixels in inter prediction. The paper proposes a low complexity learning-based inter prediction (LLIP) method to replace the traditional INTERPF. LLIP enhances the filtering process by levera
Vsevolod Evtushevsky
We describe Martin boundary of the path space of $r$-differential version of Young--Fibonacci graph. Also we establish ergodicity of the corresponding measures.
Siqi Wang, Chao Liang, Yunfan Gao, Yang Liu
Industrial parks are critical to urban economic growth. Yet, their development often encounters challenges stemming from imbalances between industrial requirements and urban services, underscoring the need for strategic planning and operations. This paper introduces IndustryScopeKG, a pioneering large-scale multi-modal, multi-level industrial park knowledge
A Study of Black Holes in $F(R)-$ModMax Gravity: Gravitational Lensing and Constraints from EHT Observations
gr-qcKhadije Jafarzade, Zeynab Bazyar, Mubasher Jamil
The study of astrophysical phenomena like black hole shadows is an effective approach to properly understand the modified gravity and explore its validity. Motivated by recent astrophysical observations, we consider a black hole (BH) in $F(R)-$ModMax gravity and study the optical features such as the shadow's geometrical shape, energy emission rate, and defl
Jaesung Kim, Jin-Hee Yoon
The near-side ridge structure has been observed in the long-range two-particle correlations in heavy-ion collisions, such as AuAu collisions at the Relativistic Heavy Ion Collider(RHIC) and PbPb collisions at the Large Hadron Collider (LHC). Hydrodynamic models have successfully explained the ridge structure in heavy-ion collisions, indicating the presence o
Tiny yet detectable WIMP-nucleon scattering cross sections in a pseudo-Nambu-Goldstone dark matter model
hep-phTomohiro Abe, Kota Ichiki
We investigate a pseudo-Nambu-Goldstone (pNG) dark matter (DM) model based on a gauged $SU(2)_x$ and a global $SU(2)_g$ symmetries. These symmetries are spontaneously broken to a global $U(1)_D$ symmetry by a vacuum expectation value of an $SU(2)_x \times SU(2)_g$ bi-fundamental scalar field. The global $SU(2)_g$ symmetry is also softly broken to a global $U
Toshiharu Chono, Hisashi Tokutomi, Kazuma Nakamura, Koji Miyazaki
We report the first spectral reflectance of tungsten carbide (WC) as potential solar selective absorber. We developed an optical measurement system for visible to mid-infrared spectroscopy, covering the range of 0.1 to 2.5 eV, to evaluate the solar selectivity. A polycrystalline WC was prepared using spark plasma sintering method. The measured spectral refle
Zihao He, Hongjie Fang, Jingjing Chen, Hao-Shu Fang
Contact-rich tasks present significant challenges for robotic manipulation policies due to the complex dynamics of contact and the need for precise control. Vision-based policies often struggle with the skill required for such tasks, as they typically lack critical contact feedback modalities like force/torque information. To address this issue, we propose F
Vsevolod Evtushevsky
For a poset $(P,\leqslant)$ we consider the first-order theory, that is defined by set $P$ and relation $\leqslant$. The problem of undecidability of combinatorial theories attracts significant attention. Recently A. Wires proved the undecidability of the elementary theory of Young lattice and also established the maximal definability property of this theory
Measurement of cross sections of $e^+e^-\to K^0_S K^0_S \psi(3686)$ from $\sqrt{s}=$ 4.682 to 4.951 GeV
hep-exBESIII Collaboration, M. Ablikim, M. N. Achasov, P. Adlarson
The process $e^+e^-\to K^0_S K^0_S \psi(3686)$ is studied by analyzing $e^+e^-$ collision data samples collected at eight center-of-mass energies ranging from 4.682 to 4.951 GeV with the BESIII detector operating at the BEPCII collider, corresponding to an integrated luminosity of $4.1~{\rm fb}^{-1}$. Observation of the $e^+e^-\to K^0_S K^0_S \psi(3686)$ pro
Chongying Dong, Feng Xu, Nina Yu
Let $V$ be a simple, rational, $C_{2}$-cofinite vertex operator algebra of CFT type, and let $k$ be a positive integer. In this paper, we determine the fusion products of twisted modules for $V^{\otimes k}$ and $G = \left\langle g \right\rangle$ generated by any permutation $g \in S_{k}$.
A Novel Approach for Bent Functions with Dillon-like Exponents and Characterizing Three Classes of Bent Functions via Kloosterman Sums
cs.DMZiran Tu, Sihem Mesnager, Xiangyong Zeng, Nian Li
Dillon-like Boolean functions are known, in the literature, to be those trace polynomial functions from $\mathbb{F}_{2^{2n}}$ to $\mathbb{F}_{2}$, with all the exponents being multiples of $2^n-1$ often called Dillon-like exponents. This paper is devoted to bent functions in which we study the bentness of some classes of Dillon-like Boolean functions connect
Enabling low threshold laser through an asymmetric tetramer metasurface harnessing polarization-independent quasi-BICs
physics.opticsT. Wang, W. Z. Di, W. E. I. Sha, R. P. Zaccaria
We propose and numerically demonstrate a novel strategy to achieve dual-band symmetry-protected bound states in the continuum (BICs) based on a nanodisk tetramer metasurface for lasing generation. The method involves breaking the in-plane symmetry along the diagonal of the metasurface unit cell by introducing air holes in the tetramers. Through our simulatio
Itinerant electron metamagnetism for lattices with van Hove density-of-states singularities near the Fermi level
cond-mat.str-elF. A. Vasilevskiy, P. A. Igoshev, V. Yu. Irkhin
Itinerant-electron metamagnetism is investigated within the Hubbard model for various lattices having van Hove singularities (vHS) in the electronic spectrum: face-centered cubic and orthorhombic lattices. The remarkable itinerant-electron metamagnetic transition occurs provided that the Fermi level is in the region with a strong positive curvature of the de
A unified variational model for grain boundary dynamics incorporating microscopic structure
cond-mat.mtrl-sciLuchan Zhang, Xiaoxue Qin, Yang Xiang
Recent experiments, atomistic simulations, and theoretical predictions have identified various new types of grain boundary motions that are controlled by the dynamics of underlying microstructure of line defects (dislocations or disconnections), to which the classical motion by mean curvature model does not apply. Different continuum models have been develop
Zhong-Yu Li, Yunheng Li, Deng-Ping Fan, Ming-Ming Cheng
Masked image modeling has achieved great success in learning representations but is limited by the huge computational costs. One cost-saving strategy makes the decoder reconstruct only a subset of masked tokens and throw the others, and we refer to this method as partial reconstruction. However, it also degrades the representation quality. Previous methods m
Conjugate Heat Transfer Effects on Bubble Growth During Flow Boiling Heat Transfer in Microchannels
physics.flu-dynOdumuyiwa A. Odumosu, Hongying Li, Tianyou Wang, Zhizhao Che
Flow boiling in microchannel heat sinks is an efficient way to dissipate high heat flux by utilizing the large surface-to-volume ratio and high latent heat. Previous studies of boiling heat transfer in microchannels mainly consider the fluid flow in channels only, but often neglect the conjugate effects of the heat conduction in the solid wall, which becomes
Yuan Liu
We prove that any nilpotent regular covering over a compact K\"ahler surface is holomorphically convex if it does not have two ends. Furthermore, we show that the Malcev covering of any compact K\"ahler manifold has at most one end.
Liran Nochumsohn, Michal Moshkovitz, Orly Avner, Dotan Di Castro
Time series forecasting is critical in numerous real-world applications, requiring accurate predictions of future values based on observed patterns. While traditional forecasting techniques work well in in-domain scenarios with ample data, they struggle when data is scarce or not available at all, motivating the emergence of zero-shot and few-shot learning s
Tavis Shore, Oscar Mendez, Simon Hadfield
Cross-view Geo-localisation is typically performed at a coarse granularity, because densely sampled satellite image patches overlap heavily. This heavy overlap would make disambiguating patches very challenging. However, by opting for sparsely sampled patches, prior work has placed an artificial upper bound on the localisation accuracy that is possible. Even
Linyi Huang, Hui Zhang, Zijian Wu, Sammy Christen
Functional grasping is essential for humans to perform specific tasks, such as grasping scissors by the finger holes to cut materials or by the blade to safely hand them over. Enabling dexterous robot hands with functional grasping capabilities is crucial for their deployment to accomplish diverse real-world tasks. Recent research in dexterous grasping, howe
Jorge Calvo-Zaragoza, Alexander Pacha, Elona Shatri
The International Workshop on Reading Music Systems (WoRMS) is a workshop that tries to connect researchers who develop systems for reading music, such as in the field of Optical Music Recognition, with other researchers and practitioners that could benefit from such systems, like librarians or musicologists. The relevant topics of interest for the workshop
LTCF-Net: A Transformer-Enhanced Dual-Channel Fourier Framework for Low-Light Image Restoration
cs.CVGaojing Zhang, Jinglun Feng
We introduce LTCF-Net, a novel network architecture designed for enhancing low-light images. Unlike Retinex-based methods, our approach utilizes two color spaces - LAB and YUV - to efficiently separate and process color information, by leveraging the separation of luminance from chromatic components in color images. In addition, our model incorporates the Tr
Di Li, Mao Yuan, Lin Wu, Jingye Yan
Long-period radio transients (LPTs) are a newly discovered class of radio emitters with periods ranging from minutes to hours. The astrophysical nature remains undetermined, particularly of LPTs with no detectable companions. We report the first evidence for a plausible supernova remnant (SNR) association with an LPT (DART J1832-0911, 2656.23+-0.15 s period)
Qifan Yu, Wei Chow, Zhongqi Yue, Kaihang Pan
Instruction-based image editing aims to modify specific image elements with natural language instructions. However, current models in this domain often struggle to accurately execute complex user instructions, as they are trained on low-quality data with limited editing types. We present AnyEdit, a comprehensive multi-modal instruction editing dataset, compr
TableTime: Reformulating Time Series Classification as Training-Free Table Understanding with Large Language Models
cs.AIJiahao Wang, Mingyue Cheng, Qingyang Mao, Yitong Zhou
Large language models (LLMs) have demonstrated their effectiveness in multivariate time series classification (MTSC). Effective adaptation of LLMs for MTSC necessitates informative data representations. Existing LLM-based methods directly encode embeddings for time series within the latent space of LLMs from scratch to align with semantic space of LLMs. Desp
Baoshun Tong, Kaiyu Song, Hanjiang Lai
Few-shot out-of-distribution (OOD) detection aims to detect OOD images from unseen classes with only a few labeled in-distribution (ID) images. To detect OOD images and classify ID samples, prior methods have been proposed by regarding the background regions of ID samples as the OOD knowledge and performing OOD regularization and ID classification optimizati
Lianghao Tan, Xiaoyi Liu, Dong Liu, Shubing Liu
To improve the convergence speed and optimization accuracy of the Dung Beetle Optimizer (DBO), this paper proposes an improved algorithm based on circle mapping and longitudinal-horizontal crossover strategy (CICRDBO). First, the Circle method is used to map the initial population to increase diversity. Second, the longitudinal-horizontal crossover strategy
Baoshun Tong, Kaiyu Song, Hanjiang Lai
Test-time adaptation with pre-trained vision-language models (VLMs) has attracted increasing attention for tackling the issue of distribution shift during the test phase. While prior methods have shown effectiveness in addressing distribution shift by adjusting classification logits, they are not optimal due to keeping text features unchanged. To address thi
Leila Gheisi, Henry Chu, Raju Gottumukkala, Yan Luo
The use of artificial intelligence (AI) in automated disease classification significantly reduces healthcare costs and improves the accessibility of services. However, this transformation has given rise to concerns about the fairness of AI, which disproportionately affects certain groups, particularly patients from underprivileged populations. Recently, a nu
Prajwal Thapa, Jinu Nyachhyon, Mridul Sharma, Bal Krishna Bal
Transformer-based pre-trained language models have dominated the field of Natural Language Processing (NLP) for quite some time now. However, the Nepali language, spoken by approximately 32 million people worldwide, remains significantly underrepresented in this domain. This underrepresentation is primarily attributed to the scarcity of monolingual data corp
Modeling of optical scattering from topographic surface measurements of high-quality mirrors
physics.opticsTomotada Akutsu, Hiroaki Yamamoto
In this paper, we revisit computational methods to obtain an angular profile of optical scattering from a smooth surface, given a two-dimensional map of topographic height errors of the surface. Quick derivations of some traditional equations and relevant references are organized to shorten the search time. A practical data-processing flow of the methods is
DynamicAvatars: Accurate Dynamic Facial Avatars Reconstruction and Precise Editing with Diffusion Models
cs.GRYangyang Qian, Yuan Sun, Yu Guo
Generating and editing dynamic 3D head avatars are crucial tasks in virtual reality and film production. However, existing methods often suffer from facial distortions, inaccurate head movements, and limited fine-grained editing capabilities. To address these challenges, we present DynamicAvatars, a dynamic model that generates photorealistic, moving 3D head
Kexin Zhang, Fuyuan Lyu, Xing Tang, Dugang Liu
The evolution of previous Click-Through Rate (CTR) models has mainly been driven by proposing complex components, whether shallow or deep, that are adept at modeling feature interactions. However, there has been less focus on improving fusion design. Instead, two naive solutions, stacked and parallel fusion, are commonly used. Both solutions rely on pre-dete
Sebastian M. Dawid, Andrew W. Jackura, Adam P. Szczepaniak
We propose a new model-independent method for determining hadronic resonances from lattice QCD. The formalism is derived from the general principles of unitarity and analyticity, as encoded in the $N/D$ representation of a partial-wave two-body amplitude. The associated quantization condition relates the finite-volume spectrum to the infinite-volume numerato
OccludeNet: A Causal Journey into Mixed-View Actor-Centric Video Action Recognition under Occlusions
cs.CVGuanyu Zhou, Wenxuan Liu, Wenxin Huang, Xuemei Jia
The lack of occlusion data in common action recognition video datasets limits model robustness and hinders consistent performance gains. We build OccludeNet, a large-scale occluded video dataset including both real and synthetic occlusion scenes in different natural settings. OccludeNet includes dynamic occlusion, static occlusion, and multi-view interactive
SURF Report: High Accuracy Methods for Computing Gravitational Potential and Gravitational Force Fields Near the Surface of Irregularly Shaped 3-Dimensional Bodies
astro-ph.EPThomas MacLean, Alan H. Barr
Accurate gravity field calculations are necessary for landing on planets, moons, asteroids, minimoons, or other irregularly shaped bodies, but current methods become increasingly inaccurate and slow near the surface. We present high accuracy, fast methods for computing gravitational potential and gravitational force fields, which are needed for future space
Dan Nissim, Danny Segev, Alfredo Torrico
The primary contribution of this paper resides in devising constant-factor approximation guarantees for revenue maximization in two-sided matching markets, under general pairwise rewards. A major distinction between our work and state-of-the-art results in this context (Ashlagi et al., 2022; Torrico et al., 2023) is that, for the first time, we are able to a
Deterministic multi-phonon entanglement between two mechanical resonators on separate substrates
quant-phMing-Han Chou, Hong Qiao, Haoxiong Yan, Gustav Andersson
Mechanical systems have emerged as a compelling platform for applications in quantum information, leveraging recent advances in the control of phonons, the quanta of mechanical vibrations. Several experiments have demonstrated control and measurement of phonon states in mechanical resonators integrated with superconducting qubits, and while entanglement of t
Occupation and diffusion of interstitial solutes in dilute alloys in perspective of the Gauss Legendre three square theorem
cond-mat.mtrl-sciXiaoshuang Wang
In the example of the diffusion of C, N, O in dilute ferric iron alloys, it is shown that the polyhedron consisting of equivalent occupation of interstitial solutes in dilute alloys, can be classified into 7 groups altogether, i.e., cube, octahedron, cuboctahedron, truncated octahedron, truncated cube, rhombicuboctahedron and truncated cuboctahedron. No more
The Visual Counter Turing Test (VCT2): A Benchmark for Evaluating AI-Generated Image Detection and the Visual AI Index (VAI)
cs.CVNasrin Imanpour, Abhilekh Borah, Shashwat Bajpai, Subhankar Ghosh
The rapid progress and widespread availability of text-to-image (T2I) generative models have heightened concerns about the misuse of AI-generated visuals, particularly in the context of misinformation campaigns. Existing AI-generated image detection (AGID) methods often overfit to known generators and falter on outputs from newer or unseen models. We introdu
Bingyao Wu, Jie-Xiang Zhu
Fix an irrational number $\alpha$. Let $X_1,X_2,\cdots$ be independent, identically distributed, integer-valued random variables with characteristic function $\varphi$, and let $S_n=\sum_{i=1}^n X_i$ be the partial sums. Consider the random walk $\{S_n \alpha\}_{n\ge 1}$ on the torus, where $\{\cdot\}$ denotes the fractional part. We study the long time asym
Optimal-rate error estimates and a twice decoupled solver for a backward Euler finite element scheme of the Doyle-Fuller-Newman model of lithium-ion cells
math.NAShu Xu, Liqun Cao
We investigate the convergence of a backward Euler finite element discretization applied to a multi-domain and multi-scale elliptic-parabolic system, derived from the Doyle-Fuller-Newman model for lithium-ion cells. We establish optimal-order error estimates for the solution in the norms $l^2(H^1)$ and $l^2(L^2(H^q_r))$, $q=0,1$. To improve computational eff
Osama A. Marzouk, E. David Huckaby
According to a recent U.S. Greenhouse Gas Emissions Inventory (1), about 42% of 2008 CO$_2$ (a greenhouse gas) emissions in the US were from burning fossil fuels (especially coal) to generate electricity. The 2010 U.S. International Energy Outlook (2) predicts that the world energy generation using coal and natural gas will continue to increase steadily in t
Research on Effectiveness Evaluation and Optimization of Baseball Teaching Method Based on Machine Learning
cs.LGShaoxuan Sun, Jingao Yuan, Yuelin Yang
In modern physical education, data-driven evaluation methods have gradually attracted attention, especially the quantitative prediction of students' sports performance through machine learning model. The purpose of this study is to use a variety of machine learning models to regress and predict students' comprehensive scores in baseball training, so as to ev
Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks
cs.CVPeng Xie, Yequan Bie, Jianda Mao, Yangqiu Song
Pre-trained vision-language models (VLMs) have showcased remarkable performance in image and natural language understanding, such as image captioning and response generation. As the practical applications of vision-language models become increasingly widespread, their potential safety and robustness issues raise concerns that adversaries may evade the system
You Li, Fan Ma, Yi Yang
The Zero-shot Composed Image Retrieval (ZSCIR) requires retrieving images that match the query image and the relative captions. Current methods focus on projecting the query image into the text feature space, subsequently combining them with features of query texts for retrieval. However, retrieving images only with the text features cannot guarantee detaile
An investigation into the performances of the Current state-of-the-art Naive Bayes, Non-Bayesian and Deep Learning Based Classifier for Phishing Detection: A Survey
cs.CRTosin Ige, Christopher Kiekintveld, Aritran Piplai, Amy Waggler
Phishing is one of the most effective ways in which cybercriminals get sensitive details such as credentials for online banking, digital wallets, state secrets, and many more from potential victims. They do this by spamming users with malicious URLs with the sole purpose of tricking them into divulging sensitive information which is later used for various cy
Denisha Thakkar, Vincent Quoc-Huy Trinh, Sonal Varma, Samira Ebrahimi Kahou
Diffusion Generative Models (DGM) have rapidly surfaced as emerging topics in the field of computer vision, garnering significant interest across a wide array of deep learning applications. Despite their high computational demand, these models are extensively utilized for their superior sample quality and robust mode coverage. While research in diffusion gen
Ziyao Zeng, Jingcheng Ni, Daniel Wang, Patrick Rim
Traditional monocular depth estimation suffers from inherent ambiguity and visual nuisances. We demonstrate that language can enhance monocular depth estimation by providing an additional condition (rather than images alone) aligned with plausible 3D scenes, thereby reducing the solution space for depth estimation. This conditional distribution is learned du
Can an increase in productivity cause a decrease in production? Insights from a model economy with AI automation
econ.GNCasey O. Barkan
It is widely assumed that increases in economic productivity necessarily lead to economic growth. In this paper, it is shown that this is not always the case. An idealized model of an economy is presented in which a new technology allows capital to be utilized autonomously without labor input. This is motivated by the possibility that advances in artificial
Matthew Pierson, Zia Mehrabi
Waterways shape earth system processes and human societies, and a better understanding of their distribution can assist in a range of applications from earth system modeling to human development and disaster response. Most efforts to date to map the world's waterways have required extensive modeling and contextual expert input, and are costly to repeat. Many
Rajiv Sambharya, Bartolomeo Stellato
We introduce a machine-learning framework to learn the hyperparameter sequence of first-order methods (e.g., the step sizes in gradient descent) to quickly solve parametric convex optimization problems. Our computational architecture amounts to running fixed-point iterations where the hyperparameters are the same across all parametric instances and consists
Wei Yuan, Guanhua Ye, Xiangyu Zhao, Quoc Viet Hung Nguyen
Time series forecasting plays a critical role in various real-world applications, including energy consumption prediction, disease transmission monitoring, and weather forecasting. Although substantial progress has been made in time series forecasting, most existing methods rely on a centralized training paradigm, where large amounts of data are collected fr
Jia yaning, Shengyong Pan
In this paper, we give an explicit formula for the rank of the $Q$-walk matrix of the Dynkin graph $A_n$. Moreover, we prove that its Smith normal form is $$ \mathrm{diag}\left( \underset{r=\lceil \frac{n}{2} \rceil}{\underbrace{1,2,2,...,2}},0,...,0 \right), $$ where $r$ is the rank of the $Q$-walk matrix $W_Q\left( A_n \right) $ of the Dynkin graph $A_n$.
Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems
cs.CEWenxiang Lin, Xinglin Pan, Shaohuai Shi, Xuan Wang
Large language models~(LLMs) are known for their high demand on computing resources and memory due to their substantial model size, which leads to inefficient inference on moderate GPU systems. Techniques like quantization or pruning can shrink model sizes but often impair accuracy, making them unsuitable for practical applications. In this work, we introduc
Yonghui Wang, Shi-Yong Chen, Zhenxing Zhou, Siyi Li
Recently, Vision Language Models (VLMs) have experienced significant advancements, yet these models still face challenges in spatial hierarchical reasoning within indoor scenes. In this study, we introduce ROOT, a VLM-based system designed to enhance the analysis of indoor scenes. Specifically, we first develop an iterative object perception algorithm using
AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks
cs.CVYou Li, Fan Ma, Yi Yang
Diffusion models have recently been employed to generate high-quality images, reducing the need for manual data collection and improving model generalization in tasks such as object detection, instance segmentation, and image perception. However, the synthetic framework is usually designed with meticulous human effort for each task due to various requirement
Yong Li
Providing optimal portfolio selection for investors has always been one of the hot topics in academia. In view of the traditional portfolio model could not adapt to the actual capital market and can provide erroneous results. This paper innovatively constructs a mean-detrended cross-correlation portfolio model (M-DCCP model), This model is designed to embed
Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation
cs.CVHaojie Zhang, Zhihao Liang, Ruibo Fu, Bingyan Liu
Long-duration talking video synthesis faces enduring challenges in achieving high video quality, portrait consistency, temporal coherence, and computational efficiency. As video length increases, issues such as visual degradation, portrait drift, temporal artifacts, and error accumulation become increasingly problematic, severely affecting the realism and re
Robustifying Long-term Human-Robot Collaboration through a Multimodal and Hierarchical Framework
cs.ROPeiqi Yu, Abulikemu Abuduweili, Ruixuan Liu, Changliu Liu
Long-term Human-Robot Collaboration (HRC) is crucial for enabling flexible manufacturing systems and integrating companion robots into daily human environments over extended periods. This paper identifies several key challenges for such collaborations, such as accurate recognition of human plan, robustness to disturbances, operational efficiency, adaptabilit
Understanding Student Acceptance, Trust, and Attitudes Toward AI-Generated Images for Educational Purposes
cs.CYAung Pyae
Recent advancements in artificial intelligence (AI) have broadened the applicability of AI-generated images across various sectors, including the creative industry and design. However, their utilization in educational contexts, particularly among undergraduate students in computer science and software engineering, remains underexplored. This study adopts an
M. Mehdi Khalighi, Doug Kelley, Jason H. Su, Brian K. Rutt
Adiabatic Bloch-Siegert B1+ mapping method addresses the long TE and high RF power deposition problems of conventional Bloch-Siegert B1+ mapping by introducing short frequency-swept ABS pulses with maximum sensitivity. Here, it is shown how maximum signal to noise ratio can be achieved in adiabatic Bloch-Siegert B1+ mapping. Signal to noise ratio of B1+ maps
Advancing Uncertain Combinatorics through Graphization, Hyperization, and Uncertainization: Fuzzy, Neutrosophic, Soft, Rough, and Beyond
cs.AITakaaki Fujita, Florentin Smarandache
Combinatorics studies how discrete objects can be counted, arranged, and combined under specified rules. Motivated by uncertainty in real-world data and decisions, modern set-theoretic formalisms such as fuzzy sets, neutrosophic sets, rough sets, soft sets, and plithogenic sets have been developed. In particular, neutrosophic sets model uncertainty by assign
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training
cs.CLXiaoye Qu, Daize Dong, Xuyang Hu, Tong Zhu
Recently, inspired by the concept of sparsity, Mixture-of-Experts (MoE) models have gained increasing popularity for scaling model size while keeping the number of activated parameters constant. In this study, we thoroughly investigate the sparsity of the dense LLaMA model by constructing MoE for both the attention (i.e., Attention MoE) and MLP (i.e., MLP Mo
Zhengyi Li, Kang Yang, Jin Tan, Wen-jie Lu
Transformer models have gained significant attention due to their power in machine learning tasks. Their extensive deployment has raised concerns about the potential leakage of sensitive information during inference. However, when being applied to Transformers, existing approaches based on secure two-party computation (2PC) bring about efficiency limitations
Jack Yu, Xueying Jia, Charlie Sun, Prince Wang
Novel view synthesis is a fundamental challenge in image-to-3D generation, requiring the generation of target view images from a set of conditioning images and their relative poses. While recent approaches like Zero-1-to-3 have demonstrated promising results using conditional latent diffusion models, they face significant challenges in generating consistent
Anisotropic anomalous diffusion in microgravity dusty plasma. Part One: Nonextensive Statistical Analysis
physics.plasm-phBradley R. Andrew, Luca Guazzotto, Lorin Matthews, Hyde Truell
Anisotropic anomalous dust diffusion in microgravity dusty plasma is investigated using experimental data from the Plasmakristall-4 (PK-4) facility on board the International Space Station. The PK-4 experiment uses video cameras to track individual dust particles, which allows the collection of large amounts of statistical information on the dust particle po
Kazuo Fujikawa, Anca Tureanu
In this note, we discuss an analogy between the BCS theory and the seesaw model of neutrinos. We believe that the analogy indicates some fundamental aspects of Majorana neutrinos. A paper on the issue has been recently presented, and we would like to describe the background of the paper together with our personal views on the problem. In essence, the convent