Skip to content

November 2024 arXiv papers — page 42

Showing 4,1014,200 of 19,800 papers

  1. Zine el abidine Kherroubi, Monika Prakash, Jean-Pierre Giacalone, Michael Baddeley

    Modern wireless communication systems have become increasingly complex due to the proliferation of wireless devices, increasing performance standards, and growing security threats. Managing these networks is becoming more challenging, requiring the use of advanced network management methods and tools. AI-driven network management systems such as Self-Optimiz

  2. I. Pavlov, A. Chaikovskaia, D. Karlovets

    We investigate the intriguing phenomenon of beta decay of a free neutron in a non-plane-wave(structured) state. Our analysis covers three types of states: unpolarized vortex (Bessel) neutrons that possess nonzero orbital angular momentum (OAM), Laguerre-Gaussian wave packets, and spin-correlated OAM (spin-orbit) states characterized by unique polarization pa

  3. Ivar Bengtsson, Anders Forsgren, Albin Fredriksson, Ye Zhang

    The steep dose gradients obtained with pencil beam scanning allow for precise tumor targeting at the cost of high sensitivity to uncertainties. Robust optimization is commonly applied to mitigate uncertainties in density and patient setup, while its application to motion management, called 4D-robust optimization (4DRO), is typically accompanied by other moti

  4. Daniela De Canditiis, Fabiano Veglianti

    The Extreme Learning Machine (ELM) is a growing statistical technique widely applied to regression problems. In essence, ELMs are single-layer neural networks where the hidden layer weights are randomly sampled from a specific distribution, while the output layer weights are learned from the data. Two of the key challenges with this approach are the architec

  5. Maurice D. Hanisch, Bence Hetényi, James R. Wootton

    Quantum error correction promises a viable path to fault-tolerant computations, enabling exponential error suppression when the device's error rates remain below the protocol's threshold. This threshold, however, strongly depends on the classical method used to decode the syndrome measurements. These classical algorithms traditionally only interpret binary d

  6. Nourelhouda Groun, Maria Villalba-Orero, Lucia Casado-Martin, Enrique Lara-Pezzi

    In the realm of cardiovascular medicine, medical imaging plays a crucial role in accurately classifying cardiac diseases and making precise diagnoses. However, the field faces significant challenges when integrating data science techniques, as a significant volume of images is required for these techniques. As a consequence, it is necessary to investigate di

  7. Yu. M. Poluektov

    In the weakly non-ideal gas model [1], the Bose-Einstein condensation at constant pressure is considered. The temperature of transition to the state with condensate is found. Temperature dependences of the total density and condensate density, the energy, entropy and heat capacities are calculated.

  8. Nicoletta Cantarini, Fabrizio Caselli, Victor Kac

    We study the embeddings of the exceptional infinite-dimensional Lie superalgebra E(1,6) in the exceptional Lie superalgebras E(5,10) and E(4,4). These questions arose in the recent works on enhanced symmetries in some supersymmetric theories by N. Garner, S. Raghavendran, I. Saberi and B. Williams.

  9. Takashi Kobayashi, Akito Noiri, Takashi Nakajima, Kenta Takeda

    An electron confined by a semiconductor quantum dot (QD) can be displaced by changes in electron occupations of surrounding QDs owing to the Coulomb interaction. For a single-spin qubit in an inhomogeneous magnetic field, such a displacement of the host electron results in a qubit energy shift which must be handled carefully for high-fidelity operations. Her

  10. A. I. Delis, N. Bekiaris-Liberis

    This work aims to provide an approach to the macroscopic modeling and simulation of pedestrian flow, coupled with contagion spreading, towards numerical investigation of the effect of certain, macro-control measures on epidemics transport dynamics. To model the dynamics of the pedestrians, a second-order macroscopic model, coupled with an Eikonal equation, i

  11. Adrien Meyer, Aditya Murali, Farahdiba Zarin, Didier Mutter

    Purpose: Automated ultrasound image analysis is challenging due to anatomical complexity and limited annotated data. To tackle this, we take a data-centric approach, assembling the largest public ultrasound segmentation dataset and training a versatile visual foundation model tailored for ultrasound. Methods: We compile US-43d, a large-scale collection of 43

  12. Yiran Wang, Willem Meijer, José Antonio Hernández López, Ulf Nilsson

    Jupyter notebooks have become central in data science, integrating code, text and output in a flexible environment. With the rise of machine learning (ML), notebooks are increasingly used for prototyping and data analysis. However, due to their dependence on complex ML libraries and the flexible notebook semantics that allow cells to be run in any order, not

  13. Hyeong-Soon Jang, Hyungjun Heo, Sangin Kim, Hyeon Hwang

    We demonstrate efficient edge couplers by fabricating a 3D mode size converter on a lithium niobate-on-insulator photonic platform. The 3D mode size converter is fabricated using an etching process that employs a Si external mask to provide height variation and adjust the width variation through tapering patterns via lithography. The measured edge coupling e

  14. Jiahui Xin, Wei Ma

    In the context of precision medicine, covariate-adjusted response-adaptive randomization (CARA) has garnered much attention from both academia and industry due to its benefits in providing ethical and tailored treatment assignments based on patients' profiles while still preserving favorable statistical properties. Recent years have seen substantial progress

  15. Manuel Knott, Divinefavour Odion, Sameer Sontakke, Anup Karwa

    Visual inspection for defect grading in agricultural supply chains is crucial but traditionally labor-intensive and error-prone. Automated computer vision methods typically require extensively annotated datasets, which are often unavailable in decentralized supply chains. We address this challenge by evaluating the Segment Anything Model (SAM) to generate de

  16. Yubin Gu, Yuan Meng, Xiaoshuai Sun, Jiayi Ji

    Multiple-in-one image restoration (IR) has made significant progress, aiming to handle all types of single degraded image restoration with a single model. However, in real-world scenarios, images often suffer from combinations of multiple degradation factors. Existing multiple-in-one IR models encounter challenges related to degradation diversity and prompt

  17. Hongdi Yang, Chengyang Li, Zhenxuan Wu, Gaozheng Li

    Soccer is a globally renowned sport with significant applications in video games and VR/AR. However, generating realistic soccer motions remains challenging due to the intricate interactions between the human player and the ball. In this paper, we introduce SMGDiff, a novel two-stage framework for generating real-time and user-controllable soccer motions. Ou

  18. Geoffrey Compère, Sk Jahanur Hoque, Emine Şeyma Kutluk

    The linear solution for quadrupolar perturbations around de Sitter spacetime was recently constructed. In this paper, we provide the flux-balance laws for each background symmetry (dilatations, rotations, spatial translations and cosmological boosts) in terms of source moments at quadrupolar order. We write the dilatation flux-balance law in two distinct way

  19. Y. Lei, N. A. Alam, Z. Z. Qin, M. Bao

    Using the charge density from the two-parameter Fermi model, a robust and nontrival correlation between binding energis and charge radii of mirror nuclei is newly proposed. This correlation enables simple yet reliable predictions of the nuclear mass and charge radius of proton-rich nuclei. The validity of these predictions is demonstrated by comparing the pr

  20. Bhuvan Sachdeva, Naren Akash, Tajamul Ashraf, Simon Mueller

    Cataract surgery is the most common surgical procedure globally, with a disproportionately higher burden in developing countries. While automated surgical video analysis has been explored in general surgery, its application to ophthalmic procedures remains limited. Existing works primarily focus on Phaco cataract surgery, an expensive technique not accessibl

  21. Jungang Li, Sicheng Tao, Yibo Yan, Xiaojie Gu

    Endeavors have been made to explore Large Language Models for video analysis (Video-LLMs), particularly in understanding and interpreting long videos. However, existing Video-LLMs still face challenges in effectively integrating the rich and diverse audio-visual information inherent in long videos, which is crucial for comprehensive understanding. This raise

  22. Marcos Marino

    In these lecture notes for the Les Houches School on Quantum Geometry I give an introductory overview of non-perturbative aspects of topological string theory. After a short summary of the perturbative aspects, I first consider the non-perturbative sectors of the theory as unveiled by the theory of resurgence. I give a self-contained derivation of recent res

  23. Alexander Zhuravlev, Yury Kurenkov, Xuchen Wang, Fedor Dushko

    One of the main applications of electromagnetic metasurfaces (MSs) is to tailor spatial field distributions. The radiation pattern of a given source can be desirably modified upon reflection on an MS having proper spatial modulation of its local macroscopic parameters. At the microscopic level, spatial modulation requires individually engineered meta-atoms a

  24. Valentin V. Gorokhovik

    In the paper we consider convex cones in infinite-dimensional real vector spaces which are endowed with no topology. The main purpose is to study an internal geometric structure of convex cones and to obtain an analytical description of those. To this end, we first introduce the notion of an open component of a convex cone and then prove that an arbitrary co

  25. Yukio Watanabe, S. Miyauchi, S. Kaku, T. Yamada

    The Work function (f)is fundamental for chemistry and electronics. Additionally, f can be used to examine the validity of the theoretical surfaces by comparing it with experimental f, even in the absence of long-range orders. In the reported and present experiments, the difference in f between pristine and oxygen-covered Au surfaces (df) is <1 eV at =<1 ML (

  26. Yuxiang Lin, Ling Luo, Ying Chen, Xushi Zhang

    Spatial transcriptomics (ST) provides high-resolution pathological images and whole-transcriptomic expression profiles at individual spots across whole-slide scales. This setting makes it an ideal data source to develop multimodal foundation models. Although recent studies attempted to fine-tune visual encoders with trainable gene encoders based on spot-leve

  27. XiaoKai Cao, WenJin Mo, ChangDong Wang, JianHuang Lai

    Vision is one of the essential sources through which humans acquire information. In this paper, we establish a novel framework for measuring image information content to evaluate the variation in information content during image transformations. Within this framework, we design a nonlinear function to calculate the neighboring information content of pixels a

  28. Bohao Chen, Yanchao Zhang, Yanan Lv, Hua Han

    Diffusion models have recently emerged as a powerful technique in image generation, especially for image super-resolution tasks. While 2D diffusion models significantly enhance the resolution of individual images, existing diffusion-based methods for 3D volume super-resolution often struggle with structure discontinuities in axial direction and high sampling

  29. Zhuoheng Li, Yaochen Wang, Zhixue Song, Yuqi Huang

    This study explores the capabilities of large language models (LLMs) in providing knowledge about cities and regions on a global scale. We employ two methods: directly querying the LLM for target variable values and extracting explicit and implicit features from the LLM correlated with the target variable. Our experiments reveal that LLMs embed a broad but v

  30. Yukti Makhija, Edward De Brouwer, Rahul G. Krishnan

    Checklists have been widely recognized as effective tools for completing complex tasks in a systematic manner. Although originally intended for use in procedural tasks, their interpretability and ease of use have led to their adoption for predictive tasks as well, including in clinical settings. However, designing checklists can be challenging, often requiri

  31. Shaohan Huang, Xun Wu, Shuming Ma, Furu Wei

    Multi-Head Mixture-of-Experts (MH-MoE) demonstrates superior performance by using the multi-head mechanism to collectively attend to information from various representation spaces within different experts. In this paper, we present a novel implementation of MH-MoE that maintains both FLOPs and parameter parity with sparse Mixture of Experts models. Experimen

  32. Estela Suarez, Hendryk Bockelmann, Norbert Eicker, Jan Eitzinger

    High-Performance Computing (HPC) systems are among the most energy-intensive scientific facilities, with electric power consumption reaching and often exceeding 20 megawatts per installation. Unlike other major scientific infrastructures such as particle accelerators or high-intensity light sources, which are few around the world, the number and size of supe

  33. Oana Boncalo, Alexandru Amaricai

    This paper proposes a new iterative gradient descent decoding method for real number parity codes. The proposed decoder, named Gradient Descent Symbol Update (GDSU), is used for a class of low-density parity-check (LDPC) real-number codes that can be defined with parity check matrices which are similar to those of the binary LDPC from communication standards

  34. Jungeun Kim, Hyeongwoo Jeon, Jongseong Bae, Ha Young Kim

    Sign language translation (SLT) is a challenging task that involves translating sign language images into spoken language. For SLT models to perform this task successfully, they must bridge the modality gap and identify subtle variations in sign language components to understand their meanings accurately. To address these challenges, we propose a novel gloss

  35. Shezheng Song, Chengxiang He, Shan Zhao, Chengyu Wang

    Multimodal large language models (MLLMs) have shown remarkable progress in high-level semantic tasks such as visual question answering, image captioning, and emotion recognition. However, despite advancements, there remains a lack of standardized benchmarks for evaluating MLLMs performance in multi-object sentiment analysis, a key task in semantic understand

  36. D. M. -A. Meyer, D. F. Torres

    In this study we quantitatively examine the manner pulsar wind, supernova ejecta and defunct stellar wind materials distribute and melt together into plerions. We performed 2.5D MHD simulations of the entire evolution of their stellar surroundings and different scenarios are explored, whether the star dies as a red supergiant and Wolf Rayet supernova progeni

  37. Hao Yi, Qingyang Li, Yulan Hu, Fuzheng Zhang

    High-quality video-text preference data is crucial for Multimodal Large Language Models (MLLMs) alignment. However, existing preference data is very scarce. Obtaining VQA preference data for preference training is costly, and manually annotating responses is highly unreliable, which could result in low-quality pairs. Meanwhile, AI-generated responses control

  38. Yuankai Liu, Lei Zhang, Jin Zhao

    The high-index saddle dynamics (HiSD) method is a powerful approach for computing saddle points and solution landscape. However, its practical applicability is constrained by the need for the explicit energy function expression. To overcome this challenge, we propose a neural network-based high-index saddle dynamics (NN-HiSD) method. It utilizes neural netwo

  39. Shuchen Weng, Haojie Zheng, Peixuan Zhang, Yuchen Hong

    We introduce VIRES, a video instance repainting method with sketch and text guidance, enabling video instance repainting, replacement, generation, and removal. Existing approaches struggle with temporal consistency and accurate alignment with the provided sketch sequence. VIRES leverages the generative priors of text-to-video models to maintain temporal cons

  40. Ruoyu Chen, Siyuan Liang, Jingzhi Li, Shiming Liu

    Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. However, interpreting these models' decisions has grown increasingly challenging. Existing interpretable attribution methods for object-level task interpretation have notable limitation

  41. Simon Müller, Ravit Helled

    The bulk-metallicity determination of giant exoplanets is essential to constrain their formation and evolution pathways and to compare them to the solar system. Previous studies inferred an inverse relation between the mass and bulk metallicity. However, the data almost exclusively contained planets that orbit FGK stars. The recent discoveries of giant exopl

  42. Yanan Wang, Zhenghao Fei, Ruichen Li, Yibin Ying

    Recent breakthroughs in large foundation models have enabled the possibility of transferring knowledge pre-trained on vast datasets to domains with limited data availability. Agriculture is one of the domains that lacks sufficient data. This study proposes a framework to train effective, domain-specific, small models from foundation models without manual ann

  43. Giovanni Barbarino, Nicolas Gillis

    The successive projection algorithm (SPA) is a workhorse algorithm to learn the $r$ vertices of the convex hull of a set of $(r-1)$-dimensional data points, a.k.a. a latent simplex, which has numerous applications in data science. In this paper, we revisit the robustness to noise of SPA and several of its variants. In particular, when $r \geq 3$, we prove th

  44. KaiZhou Li, Jindong Gu, Xinchun Yu, Junjie Cao

    The security risks of AI-driven video editing have garnered significant attention. Although recent studies indicate that adding perturbations to images can protect them from malicious edits, directly applying image-based methods to perturb each frame in a video becomes ineffective, as video editing techniques leverage the consistency of inter-frame informati

  45. E. O. Pozdeeva

    In the Einstein--Gauss--Bonnet (EGB) gravity models, the slow-roll approximation has been extended by taking into account the first-order slow-roll parameter $\delta_1 =-2\,H^2\,\xi^\prime/U_0$, which is proportional to the first derivative of the Gauss-Bonnet coupling function $\xi$ with respect to the e-folding number. These extensions lead to the question

  46. Aishwarya Agarwal, Srikrishna Karanam, Vineet Gandhi

    We consider the problem of single-source domain generalization. Existing methods typically rely on extensive augmentations to synthetically cover diverse domains during training. However, they struggle with semantic shifts (e.g., background and viewpoint changes), as they often learn global features instead of local concepts that tend to be domain invariant.

  47. Dong Chen

    In the era of AI, recommendation algorithms and generative AI challenge information autonomy by creating echo chambers and blurring the line between authentic and fabricated content. The Critical Canvas addresses these challenges with a novel information exploration platform designed to restore balance between algorithmic efficiency and human agency. It empl

  48. Yongchang Hui, Yuteng Zhang, Siting Huang

    This article considers to model large-dimensional matrix time series by introducing a regression term to the matrix factor model. This is an extension of classic matrix factor model to incorporate the information of known factors or useful covariates. We establish the convergence rates of coefficient matrix, loading matrices and the signal part. The theoreti

  49. Alemayehu Nana Koya, Anastasiia Sapunova, Nageswar Reddy Sanamreddy, Yanqiu Zou

    The compositional asymmetry of Janus micro- and nanoparticles gives unprecedented opportunities to manipulate such composite particles with different stimuli to achieve enhanced optical, magnetic and photothermal responses, which can be exploited for sensing, phototherapy, and nanoscale robotic applications. This perspective overviews recent advances in opti

  50. Liangliang Huang, Xiangang Wan, Feng Tang

    The calculation of (co)irreducible representations of energy bands at high-symmetry points (HSPs) is essential for high-throughput research on topological materials based on symmetry-indicators or topological quantum chemistry. However, existing computational packages usually require transforming crystal structures into specific conventions, thus hindering e

  51. Zhihua Duan, Jialin Wang

    Large Language Models (LLMs) still face challenges when dealing with complex reasoning tasks, often resulting in hallucinations, which limit the practical application of LLMs. To alleviate this issue, this paper proposes a new method that integrates different LLMs to expand the knowledge boundary, reduce dependence on a single model, and promote in-depth deb

  52. Bo-Wen Yu, Bang-Gui Liu

    Altermagnetic phase is recently found as a new magnetic phase in addition to the conventional collinear spin orders, and great efforts have been made to explore novel effects and potential applications in such materials. Here, we show that there are robust altermagnetic spin-split flat bands near the valence band edge in rutile CoF$_2$ through first-principl

  53. Zhe Wang, Nan Li, Yansha Deng, A. Hamid Aghvami

    The emergence of the metaverse has boosted productivity and creativity, driving real-time updates and personalized content, which will substantially increase data traffic. However, current bit-oriented communication networks struggle to manage this high volume of dynamic information, restricting metaverse applications interactivity. To address this research

  54. Wei Ai, Jianbin Li, Ze Wang, Yingying Wei

    Graph contrastive learning has been successfully applied in text classification due to its remarkable ability for self-supervised node representation learning. However, explicit graph augmentations may lead to a loss of semantics in the contrastive views. Secondly, existing methods tend to overlook edge features and the varying significance of node features

  55. Jiajun Luo, Lizhuo Luo, Jianru Xu, Jiajun Song

    Mixture-of-Experts-based (MoE-based) diffusion models demonstrate remarkable scalability in high-fidelity image generation, yet their reliance on expert parallelism introduces critical communication bottlenecks. State-of-the-art methods alleviate such overhead in parallel diffusion inference through computation-communication overlapping, termed displaced par

  56. Masaya Arahata, Shota Kita, Kazuo Aoyama, Akihiko Shinya

    Optical neural network (ONN) has been attracting intense attention owing to their low latency and low-power consumption. Among the ONNs, optical recurrent neural network (RNN) enables low-power and high-speed time-series data processing using a compact loop structure. The loop losses need to be efficiently compensated so that the time-series information is m

  57. Vladimir Yugay, Theo Gevers, Martin R. Oswald

    Simultaneous localization and mapping (SLAM) systems with novel view synthesis capabilities are widely used in computer vision, with applications in augmented reality, robotics, and autonomous driving. However, existing approaches are limited to single-agent operation. Recent work has addressed this problem using a distributed neural scene representation. Un

  58. Qiao Yu, Xianzhi Li, Yuan Tang, Xu Han

    Generating 3D meshes from a single image is an important but ill-posed task. Existing methods mainly adopt 2D multiview diffusion models to generate intermediate multiview images, and use the Large Reconstruction Model (LRM) to create the final meshes. However, the multiview images exhibit local inconsistencies, and the meshes often lack fidelity to the inpu

  59. Ryutaro Tsuji, Yasumichi Aoki, Ken-Ichi Ishikawa, Yoshinobu Kuramashi

    We present the results for the nucleon axial-vector, induced pseudoscalar and pion-nucleon couplings obtained from 2+1 flavor lattice QCD at the physical point with a large spatial extent of about 10 fm. Our calculations are performed with the PACS10 gauge configurations generated by the PACS Collaboration with the six stout-smeared $O(a)$ improved Wilson-cl

  60. Anumita Bose, Shubham Purwar, Setti Thirupathaiah, Awadhesh Narayan

    Recently, time-reversal symmetry broken magnetic Weyl semimetals (WSMs) have attracted extensive attention and have provided an intriguing platform for exploring fundamental physical phenomena. The study of chromium telluride-based systems has also drawn significant interest towards spintronics applications owing to their high Curie temperatures. Here, using

  61. Phuc Nguyen, Minh Luu, Anh Tran, Cuong Pham

    Existing 3D instance segmentation methods frequently encounter issues with over-segmentation, leading to redundant and inaccurate 3D proposals that complicate downstream tasks. This challenge arises from their unsupervised merging approach, where dense 2D instance masks are lifted across frames into point clouds to form 3D candidate proposals without direct

  62. Andrey L. Delitsyn, Irina K. Troshina

    It has been proven that when connecting two infinite semi-cylinders or waveguides with a finite cylinder or resonator at a certain frequency, it is possible to transmit a signal almost completely from one semi-cylinder to another. In this case, the reflected field is arbitrarily small. A very simple technique based on the expansion of the solution in a Fouri

  63. A. Aynbund, V. V. Kiselev

    We study a spatial-temporal structure of quantum fluctuations in the stress-energy tensor of zero-point modes for a scalar field in order to formulate a covariant model. The model describes an invariant vacuum contribution to the cosmological constant in the non-stationary coherent state in a finite volume. Bare and effective mean values of vacuum energy den

  64. Wenhao Xu, Wenming Weng, Yueyi Zhang, Ruikang Xu

    Deformable 3D Gaussian Splatting (3D-GS) is limited by missing intermediate motion information due to the low temporal resolution of RGB cameras. To address this, we introduce the first approach combining event cameras, which capture high-temporal-resolution, continuous motion data, with deformable 3D-GS for dynamic scene reconstruction. We observe that thre

  65. Mads Hustad Sandøy

    We classify self-injective radical cube zero algebras with respect to whether they satisfy certain finite generation conditions sufficient to have a fruitful theory of support varieties defined via Hochschild cohomology in the vein of (Erdmann et al, 2004) and (Snashall and Solberg, 2004). Using skew group algebras and Linckelmann's notion of separable equiv

  66. Aravindan Sundaram, Ujjayan Pal, Abhimanyu Chauhan, Aishwarya Agarwal

    Despite recent advancements in text-to-image models, achieving semantically accurate images in text-to-image diffusion models is a persistent challenge. While existing initial latent optimization methods have demonstrated impressive performance, we identify two key limitations: (a) attention neglect, where the synthesized image omits certain subjects from th

  67. Thomas Gauthier, Gabriel Vigny

    We prove that several dynamically defined fractals in $\mathbb{C}$ and $\mathbb{C}^2$ which arise from different type of polynomial dynamical systems can not be the same objects. One of our main results is that the closure of Misiurewicz PCF cubic polynomials (the strong bifurcation locus) cannot be the Julia set of a regular polynomial endomorphism of $\mat

  68. Holger Dette, Marius Kroll

    We take a different look at the problem of testing the independence of two metric-space-valued random variables using the distance correlation. Instead of testing if the distance correlation vanishes exactly, we are interested in the hypothesis that it does not exceed a certain threshold. Our testing problem is motivated by the observation that in many cases

  69. Niklas Beisert, Benedikt König

    In this article, we reconsider the formulation of Yangian symmetry for planar N=4 supersymmetric Yang-Mills theory, and we investigate to what extent this symmetry lifts to the beta/gamma-deformation of the model. We first apply cohomology of variational forms towards a thorough derivation of the invariance statement for the undeformed action from covariance

  70. Chuan Liu, Huanran Chen, Yichi Zhang, Jun Zhu

    Adversarial examples exhibit cross-model transferability, enabling threatening black-box attacks on commercial models. Model ensembling, which attacks multiple surrogate models, is a known strategy to improve this transferability. However, prior studies typically use small, fixed ensembles, which leaves open an intriguing question of whether scaling the numb

  71. Yuehan Zhang, Angela Yao

    Self-supervised learning is crucial for super-resolution because ground-truth images are usually unavailable for real-world settings. Existing methods derive self-supervision from low-resolution images by creating pseudo-pairs or by enforcing a low-resolution reconstruction objective. These methods struggle with insufficient modeling of real-world degradatio

  72. Yiheng Li, Ruibing Hou, Hong Chang, Shiguang Shan

    Human pose plays a crucial role in the digital age. While recent works have achieved impressive progress in understanding and generating human poses, they often support only a single modality of control signals and operate in isolation, limiting their application in real-world scenarios. This paper presents UniPose, a framework employing Large Language Model

  73. Shu-Xu Yi, Emre Seyit Yorgancioglu, S. -L. Xiong, S. -N. Zhang

    Recent observations have challenged the long-held opinion that the duration of gamma-ray burst (GRB) prompt emission is determined by the activity epochs of the central engine. Specifically, the observations of GRB 230307A have revealed a different scenario in which the duration of the prompt emission is predominantly governed by the energy dissipation proce

  74. Junho Kim, Hyunjun Kim, Hosu Lee, Yong Man Ro

    Despite advances in Large Multi-modal Models, applying them to long and untrimmed video content remains challenging due to limitations in context length and substantial memory overhead. These constraints often lead to significant information loss and reduced relevance in the model responses. With the exponential growth of video data across web platforms, und

  75. Vinayak Gupta, Manoj S, Mukund Varma T, Kaushik Mitra

    Underwater images suffer from colour shifts, low contrast, and haziness due to light absorption, refraction, scattering and restoring these images has warranted much attention. In this work, we present Unsupervised Underwater Neural Radiance Field U2NeRF, a transformer-based architecture that learns to render and restore novel views conditioned on multi-view

  76. Mischa Dombrowski, Weitong Zhang, Sarah Cechnicka, Hadrien Reynaud

    Generative methods now produce outputs nearly indistinguishable from real data but often fail to fully capture the data distribution. Unlike quality issues, diversity limitations in generative models are hard to detect visually, requiring specific metrics for assessment. In this paper, we draw attention to the current lack of diversity in generative models a

  77. Sayani Maity, Prabir Rudra

    In this paper, we have studied the effects of holographic dark energy on the evolution of gravitational waves. The background evolution of gravitational waves in a flat FRW universe is considered and studied in the presence of various holographic dark energy models. The perturbation equations governing the evolution of the gravitational waves have been const

  78. Jinpeng Liu, Jiale Xu, Weihao Cheng, Yiming Gao

    We introduce NovelGS, a diffusion model for Gaussian Splatting (GS) given sparse-view images. Recent works leverage feed-forward networks to generate pixel-aligned Gaussians, which could be fast rendered. Unfortunately, the method was unable to produce satisfactory results for areas not covered by the input images due to the formulation of these methods. In

  79. Yuan Zhou, Qingshan Xu, Jiequan Cui, Junbao Zhou

    Recently, large efforts have been made to design efficient linear-complexity visual Transformers. However, current linear attention models are generally unsuitable to be deployed in resource-constrained mobile devices, due to suffering from either few efficiency gains or significant accuracy drops. In this paper, we propose a new de\textbf{C}oupled du\textbf

  80. Wang Yu, Wei Wei

    Recognition of low-quality face images remains a challenge due to invisible or deformation in partial facial regions. For low-quality images dominated by missing partial facial regions, local region similarity contributes more to face recognition (FR). Conversely, in cases dominated by local face deformation, excessive attention to local regions may lead to

  81. Mahela Pandukabhaya, Tharaka Fonseka, Madhumini Kulathunge, Roshan Godaliyadda

    Mastering psychomotor skills, such as those essential in sports, rehabilitation, and professional training, often requires a precise understanding of motion patterns and performance metrics. This study proposes a versatile framework for optimizing psychomotor learning through human motion analysis. Utilizing a wearable IMU sensor system, the motion trajector

  82. Jiarui Song, Yingbo Sun, Qing Dong, Xuewu Ji

    This paper studies the trajectory tracking and motion control problems for autonomous vehicles (AVs). A parameter adaptive control framework for AVs is proposed to enhance tracking accuracy and yaw stability. While establishing linear quadratic regulator (LQR) and three robust controllers, the control framework addresses trajectory tracking and motion contro

  83. Ion Stoica, Matei Zaharia, Joseph Gonzalez, Ken Goldberg

    Despite the significant strides made by generative AI in just a few short years, its future progress is constrained by the challenge of building modular and robust systems. This capability has been a cornerstone of past technological revolutions, which relied on combining components to create increasingly sophisticated and reliable systems. Cars, airplanes,

  84. Xingshuo Han, Xuanye Zhang, Xiang Lan, Haozhao Wang

    By using a control variate to calibrate the local gradient of each client, Scaffold has been widely known as a powerful solution to mitigate the impact of data heterogeneity in Federated Learning. Although Scaffold achieves significant performance improvements, we show that this superiority is at the cost of increased security vulnerabilities. Specifically,

  85. Ian DSouza, Chris Gordon, John C. Forbes

    If QCD axion dark matter formed post-inflation, axion miniclusters emerged from isocurvature fluctuations and later merged hierarchically into minihalos. These minihalos, potentially disrupted by stellar encounters in the Milky Way, affect axion detectability. We extend prior analyses by more accurately incorporating multiple stellar encounters and dynamical

  86. Changqing JI

    In the application of brain-computer interface (BCI), we not only need to accurately decode brain signals,but also need to consider the explainability of the decoding process, which is related to the reliability of the model. In the process of designing a decoder or processing brain signals, we need to explain the discovered phenomena in physical or physiolo

  87. Nonghai Zhang, Hao Tang

    When humans read a specific text, they often visualize the corresponding images, and we hope that computers can do the same. Text-to-image synthesis (T2I), which focuses on generating high-quality images from textual descriptions, has become a significant aspect of Artificial Intelligence Generated Content (AIGC) and a transformative direction in artificial

  88. Yukun Zhang, Yusen Wu, Xiao Yuan

    The local Hamiltonian (LH) problem, the quantum analog of the classical constraint satisfaction problem, is a cornerstone of quantum computation and complexity theory. It is known to be QMA-complete, indicating that it is challenging even for quantum computers. Interestingly, the guided local Hamiltonian (GLH) problem -- an LH problem with a guiding state th

  89. Bo Liu, Ke Zou, Liming Zhan, Zexin Lu

    Medical Visual Question Answering (Med-VQA) combines computer vision and natural language processing to automatically answer clinical inquiries about medical images. However, current Med-VQA datasets exhibit two significant limitations: (1) they often lack visual and textual explanations for answers, hindering comprehension for patients and junior doctors; (

  90. Yaniv Nemcovsky, Avi Mendelson, Chaim Baskin

    Sparse and patch adversarial attacks were previously shown to be applicable in realistic settings and are considered a security risk to autonomous systems. Sparse adversarial perturbations constitute a setting in which the adversarial perturbations are limited to affecting a relatively small number of points in the input. Patch adversarial attacks denote the

  91. Sunghwan Kim, Kwangwook Seo, Tongyoung Kim, Jinyoung Yeo

    Recent approaches in Conversational Recommender Systems (CRSs) have tried to simulate real-world users engaging in conversations with CRSs to create more realistic testing environments that reflect the complexity of human-agent dialogue. Despite the significant advancements, reliably evaluating the capability of CRSs to elicit user preferences still faces a

  92. Kangao Ouyang, Fengxian Tang, Zhilin Yuan, Jun Li

    Traditional standard single-mode fibers (SSMF) are unable to satisfy the future long-distance and high-speed optical channel transmission requirement due to their relatively large signal losses. To address this issue, the ultra-low loss and large effective area (ULL) fibers are successfully manufactured and expected to deployed in the existing optical networ

  93. Yu Zhang, Mingzi Wang, Lancheng Zou, Wulong Liu

    Transformer-based large language models (LLMs) have achieved remarkable success as model sizes continue to grow, yet their deployment remains challenging due to significant computational and memory demands. Quantization has emerged as a promising solution, and state-of-the-art quantization algorithms for LLMs introduce the need for mixed-precision matrix mul

  94. Chenjie Cao, Chaohui Yu, Shang Liu, Fan Wang

    We introduce MVGenMaster, a multi-view diffusion model enhanced with 3D priors to address versatile Novel View Synthesis (NVS) tasks. MVGenMaster leverages 3D priors that are warped using metric depth and camera poses, significantly enhancing both generalization and 3D consistency in NVS. Our model features a simple yet effective pipeline that can generate u

  95. Yicheng Feng, Yijiang Li, Wanpeng Zhang, Hao Luo

    We present VideoOrion, a Video Large Language Model (Video-LLM) that explicitly captures the key semantic information in videos - the spatial-temporal dynamics of objects throughout the videos. VideoOrion employs expert vision models to extract object dynamics through a detect-segment-track pipeline, encoding them into a set of object tokens by aggregating s

  96. Jiawei Li, Ka-di Zhu

    The symmetron, one of the light scalar fields introduced by dark energy theories, is thought to modify the gravitational force when it couples to matter. However, detecting the symmetron field is challenging due to its screening behavior in the high-density environment of traditional measurements. In this paper, we propose a scheme to set constraints on the

  97. Toyotaro Suzumura, Hiroki Kanezashi, Shotaro Akahori

    In diagnosing neurological disorders from electroencephalography (EEG) data, foundation models such as Transformers have been employed to capture temporal dynamics. Additionally, Graph Neural Networks (GNNs) are critical for representing the spatial relationships among EEG sensors. However, fine-tuning these large-scale models for both temporal and spatial f

  98. Sizai Hou, Songze Li, Duanyi Yao

    Self-supervised learning (SSL) is pervasively exploited in training high-quality upstream encoders with a large amount of unlabeled data. However, it is found to be susceptible to backdoor attacks merely via polluting a small portion of training data. The victim encoders associate triggered inputs with target embeddings, e.g., mapping a triggered cat image t

  99. C. A. Avellaneda, O. O. Melo, N. A. Cruz

    This paper proposes an analysis methodology for the case where there is longitudinal data with destructive sampling of observational units, which come from experimental units that are measured at all times of the analysis. A mixed linear model is proposed and compared with regression models with fixed and mixed effects, among which is a similar that is used

  100. Harsh Goel, Sai Shankar Narasimhan, Oguzhan Akcin, Sandeep Chinchali

    In recent years, significant progress has been made in collecting large-scale datasets to improve segmentation and autonomous driving models. These large-scale datasets are often dominated by common environmental conditions such as "Clear and Day" weather, leading to decreased performance in under-represented conditions like "Rainy and Night". To address thi