Skip to content

March 2024 arXiv papers — page 95

Showing 9,4019,500 of 20,618 papers

  1. Mengwei Wang, Ruixin Yan, Zeyi Hou, Ning Lang

    In the field of medical image analysis, the scarcity of Chinese chest X-ray report datasets has hindered the development of technology for generating Chinese chest X-ray reports. On one hand, the construction of a Chinese chest X-ray report dataset is limited by the time-consuming and costly process of accurate expert disease annotation. On the other hand, a

  2. Yue Ding, Hongqiao Shi, Shuang Song, Yonghui Wang

    The integration of local elements into shape contours is critical for target detection and identification in cluttered scenes. Previous studies have shown that observers can learn to use image regularities for contour integration and target identification. However, we still know little about the generalization of perceptual learning in contour integration. S

  3. Amira Guesmi, Muhammad Abdullah Hanif, Ihsen Alouani, Bassem Ouni

    Monocular depth estimation (MDE) has advanced significantly, primarily through the integration of convolutional neural networks (CNNs) and more recently, Transformers. However, concerns about their susceptibility to adversarial attacks have emerged, especially in safety-critical domains like autonomous driving and robotic navigation. Existing approaches for

  4. Tobias Stollenwerk, Stuart Hadfield

    Parameterized quantum circuits are attractive candidates for potential quantum advantage in the near term and beyond. At the same time, as quantum computing hardware not only continues to improve but also begins to incorporate new features such as mid-circuit measurement and adaptive control, opportunities arise for innovative algorithmic paradigms. In this

  5. Joonhyung Lee, Sangbeom Park, Yongin Kwon, Jemin Lee

    In robotic object manipulation, human preferences can often be influenced by the visual attributes of objects, such as color and shape. These properties play a crucial role in operating a robot to interact with objects and align with human intention. In this paper, we focus on the problem of inferring underlying human preferences from a sequence of raw visua

  6. Yizheng Wang, Xiang Li, Ziming Yan, Shuaifeng Ma

    Homogenization is a fundamental tool for studying multiscale physical phenomena. Traditional numerical homogenization methods, heavily reliant on finite element analysis, demand significant computational resources, especially for complex geometries, materials, and high-resolution problems. To address these challenges, we propose PreFine-Homo, a novel numeric

  7. Hyoungjun Kim, Sungjong No, Hyungkee Yoo

    The linking number of an oriented two-component link is an invariant indicating how intertwined the two components are. Tuler proved that the linking number of a two-component rational $\frac{p}{q}$-link is $$\sum^{\frac{|p|}{2}}_{k=1} (-1)^{\big\lfloor (2k-1) \frac{q}{p} \big\rfloor }.$$ In this paper, we provide a simple proof the above result, and introdu

  8. Haoxiang Ma, Ran Qin, Modi shi, Boyang Gao

    This paper focuses on the sim-to-real issue of RGB-D grasp detection and formulates it as a domain adaptation problem. In this case, we present a global-to-local method to address hybrid domain gaps in RGB and depth data and insufficient multi-modal feature alignment. First, a self-supervised rotation pre-training strategy is adopted to deliver robust initia

  9. Sungphill Moon, Hyeontae Son, Dongcheol Hur, Sangwook Kim

    Despite the progress of learning-based methods for 6D object pose estimation, the trade-off between accuracy and scalability for novel objects still exists. Specifically, previous methods for novel objects do not make good use of the target object's 3D shape information since they focus on generalization by processing the shape indirectly, making them less e

  10. Shenyu Zhang, Yu Li, Rui Wu, Xiutian Huang

    Automatic methods for evaluating machine-generated texts hold significant importance due to the expanding applications of generative systems. Conventional methods tend to grapple with a lack of explainability, issuing a solitary numerical score to signify the assessment outcome. Recent advancements have sought to mitigate this limitation by incorporating lar

  11. Takuya Fujimura, Keisuke Imoto, Tomoki Toda

    We propose discriminative neighborhood smoothing of generative anomaly scores for anomalous sound detection. While the discriminative approach is known to achieve better performance than generative approaches often, we have found that it sometimes causes significant performance degradation due to the discrepancy between the training and test data, making it

  12. Juming Xiong, Ethan H. Nguyen, Yilin Liu, Ruining Deng

    Recently, circle representation has been introduced for medical imaging, designed specifically to enhance the detection of instance objects that are spherically shaped (e.g., cells, glomeruli, and nuclei). Given its outstanding effectiveness in instance detection, it is compelling to consider the application of circle representation for segmenting instance m

  13. Dazhao Du, Enhan Li, Lingyu Si, Fanjiang Xu

    Underwater video enhancement (UVE) aims to improve the visibility and frame quality of underwater videos, which has significant implications for marine research and exploration. However, existing methods primarily focus on developing image enhancement algorithms to enhance each frame independently. There is a lack of supervised datasets and models specifical

  14. Ramy Farag, Parth Upadhyay, Yixiang Gao, Jacket Demby

    Manual analysis and diagnosis of COVID-19 through the examination of Computed Tomography (CT) images of the lungs can be time-consuming and result in errors, especially given high volume of patients and numerous images per patient. So, we address the need for automation of this task by developing a new deep learning model-based pipeline. Our motivation was s

  15. Azad Singh, Vandan Gorade, Deepak Mishra

    Self-supervised learning (SSL) is potentially useful in reducing the need for manual annotation and making deep learning models accessible for medical image analysis tasks. By leveraging the representations learned from unlabeled data, self-supervised models perform well on tasks that require little to no fine-tuning. However, for medical images, like chest

  16. Ruicheng Wang, Jianfeng Xiang, Jiaolong Yang, Xin Tong

    We propose a novel image editing technique that enables 3D manipulations on single images, such as object rotation and translation. Existing 3D-aware image editing approaches typically rely on synthetic multi-view datasets for training specialized models, thus constraining their effectiveness on open-domain images featuring significantly more varied layouts

  17. Jiasheng Wu, Shaojie Su, Xiong Wang, Jingjing Zhang

    The construction of Low Earth Orbit (LEO) satellite constellations has recently spurred tremendous attention from academia and industry. 5G and 6G standards have specified LEO satellite network as a key component of 5G and 6G networks. However, ground terminals experience frequent, high-latency handover incurred by satellites' fast travelling speed, which de

  18. Taner Akgün, Clàudia Soriano-Guerrero, Albert Elias-López, Daniele Viganò

    The inflated radii observed in hundreds of Hot Jupiters represent a long-standing open issue. The observed correlation between radii and irradiation strength, and the occasional extreme cases, nearly double the size of Jupiter, remain without a comprehensive quantitative explanation. In this investigation, we delve into this issue within the framework of Ohm

  19. Florian Schweiger, Wei Wu, Ofer Zeitouni

    We consider the discrete Ginzburg-Landau field with potential satisfying a uniform convexity condition, in the critical dimension $d=2$, and prove that its maximum over boxes of sidelength $N$, centered by an explicit $N$-dependent centering, is tight.

  20. Shek Yeung, Wangzheng Zhang, Ming-chung Chu

    A simple and natural extension of the standard Lambda cold dark matter ($\Lambda$CDM) model is to allow relic neutrinos to have finite chemical potentials. We confront this $\Lambda$CDM$\xi$ model, a $\Lambda$CDM with neutrino mass $M_\nu$ and degeneracy $\xi_3$ as additional parameters, with various cosmological data sets. We find that the $H_0$ and $S_8$ t

  21. Runtian Yuan, Qingqiu Li, Junlin Hou, Jilan Xu

    In response to the need for rapid and accurate COVID-19 diagnosis during the global pandemic, we present a two-stage framework that leverages pseudo labels for domain adaptation to enhance the detection of COVID-19 from CT scans. By utilizing annotated data from one domain and non-annotated data from another, the model overcomes the challenge of data scarcit

  22. Qizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt

    Large vision language models, such as CLIP, demonstrate impressive robustness to spurious features than single-modal models trained on ImageNet. However, existing test datasets are typically curated based on ImageNet-trained models, which aim to capture the spurious features inherited in ImageNet. Benchmarking CLIP models based on the ImageNet-oriented spuri

  23. Thien-Minh Nguyen, Shenghai Yuan, Thien Hoang Nguyen, Pengyu Yin

    Perception plays a crucial role in various robot applications. However, existing well-annotated datasets are biased towards autonomous driving scenarios, while unlabelled SLAM datasets are quickly over-fitted, and often lack environment and domain variations. To expand the frontier of these fields, we introduce a comprehensive dataset named MCD (Multi-Campus

  24. Yile Chen, Xiucheng Li, Gao Cong, Zhifeng Bao

    In this study, we introduce a novel framework called Toast for learning general-purpose representations of road networks, along with its advanced counterpart DyToast, designed to enhance the integration of temporal dynamics to boost the performance of various time-sensitive downstream tasks. Specifically, we propose to encode two pivotal semantic characteris

  25. Mrityunjoy Gain, Avi Deb Raha, Rameswar Debnath

    In this paper, we formulate the colorization problem into a multinomial classification problem and then apply a weighted function to classes. We propose a set of formulas to transform color values into color classes and vice versa. To optimize the classes, we experiment with different bin sizes for color class transformation. Observing class appearance, stan

  26. Kanchan Mittal, Pankaj Gautam, V. Vetrivel

    We introduce a forward-backward-forward (FBF) algorithm for solving bilevel equilibrium problem associated with bifunctions on a real Hilbert space. This modifies the forward-backward algorithm by relaxing cocoercivity with monotone and Lipschitzness. Further, we present the FBF dynamical system and investigate the generated trajectory's existence, uniquenes

  27. Yang Zhou, Hao Shao, Letian Wang, Steven L. Waslander

    Predicting the future motion of surrounding agents is essential for autonomous vehicles (AVs) to operate safely in dynamic, human-robot-mixed environments. Context information, such as road maps and surrounding agents' states, provides crucial geometric and semantic information for motion behavior prediction. To this end, recent works explore two-stage predi

  28. Mingkui Tan, Guohao Chen, Jiaxiang Wu, Yifan Zhang

    Test-time adaptation (TTA) seeks to tackle potential distribution shifts between training and test data by adapting a given model w.r.t. any test sample. Although recent TTA has shown promising performance, we still face two key challenges: 1) prior methods perform backpropagation for each test sample, resulting in unbearable optimization costs to many appli

  29. Tohru Mashiko, Tsuyoshi Okubo

    Besides the exactly solvable spin-1/2 Kitaev model, higher spin-$S$ ones, not exactly solvable, are promising playgrounds for researches on the quantum spin liquid as well. One of the main interests in higher spin-S cases is the interplay between the Kitaev spin liquid (KSL) and spin nematics. We probe this interplay in a spin-1 model on the honeycomb lattic

  30. Holger Dette, Jiajun Tang

    For statistical inference on an infinite-dimensional Hilbert space $\H $ with no moment conditions we introduce a new class of energy distances on the space of probability measures on $\H$. The proposed distances consist of the integrated squared modulus of the corresponding difference of the characteristic functionals with respect to a reference probability

  31. Keisuke Izumi, Toshifumi Noumi, Daisuke Yoshida

    We investigate a gedanken experiment to destroy an extremally charged black hole by dropping a test particle, provided that there are multiple $U(1)$ gauge fields coupled with each other through higher derivative interactions. In the absence of higher derivative corrections, it is known that the Coulomb repulsion prevents a test particle that would break the

  32. Vishnu Sashank Dorbala, Sanjoy Chowdhury, Dinesh Manocha

    We present a novel approach to automatically synthesize "wayfinding instructions" for an embodied robot agent. In contrast to prior approaches that are heavily reliant on human-annotated datasets designed exclusively for specific simulation platforms, our algorithm uses in-context learning to condition an LLM to generate instructions using just a few referen

  33. Yanchang Fu, Pei Xu, Dongdong Bai, Lingyun Zhao

    Hand abstraction has been instrumental in developing powerful AI for Texas Hold'em poker, a widely studied testbed for imperfect information games (IIGs). Despite its success, the hand abstraction task lacks robust theoretical tools, limiting both algorithmic innovation and theoretical progress. To address this, we extend the IIG framework with the \textbf{s

  34. Farnaz Jahanbakhsh, David R. Karger

    The status-quo of misinformation moderation is a central authority, usually social platforms, deciding what content constitutes misinformation and how it should be handled. However, to preserve users' autonomy, researchers have explored democratized misinformation moderation. One proposition is to enable users to assess content accuracy and specify whose ass

  35. Xinrun Xu, Manying Lv, Zhanbiao Lian, Yurong Wu

    The clustering method based on graph models has garnered increased attention for its widespread applicability across various knowledge domains. Its adaptability to integrate seamlessly with other relevant applications endows the graph model-based clustering analysis with the ability to robustly extract "natural associations" or "graph structures" within data

  36. Kai Chen, Haichao Liu, Yulin Li, Jianghua Duan

    Compared to conventional decomposition methods that use ellipses or polygons to represent free space, starshaped representation can better capture the natural distribution of sensor data, thereby exploiting a larger portion of traversable space. This paper introduces a novel motion planning and control framework for navigating robots in unknown and cluttered

  37. Yanling Wang, Jing Zhang, Lingxi Zhang, Lixin Liu

    Open-world semi-supervised learning (Open-world SSL) for node classification, that classifies unlabeled nodes into seen classes or multiple novel classes, is a practical but under-explored problem in the graph community. As only seen classes have human labels, they are usually better learned than novel classes, and thus exhibit smaller intra-class variances

  38. Shuang Wang, Fei Deng, Peifan Jiang, Zishan Gong

    Geographical, physical, or economic constraints often result in missing traces within seismic data, making the reconstruction of complete seismic data a crucial step in seismic data processing. Traditional methods for seismic data reconstruction require the selection of multiple empirical parameters and struggle to handle large-scale continuous missing data.

  39. Yue Fan, Xiaojian Ma, Rujie Wu, Yuntao Du

    We explore how reconciling several foundation models (large language models and vision-language models) with a novel unified memory mechanism could tackle the challenging video understanding problem, especially capturing the long-term temporal relations in lengthy videos. In particular, the proposed multimodal agent VideoAgent: 1) constructs a structured mem

  40. Debanjali Bhattacharya, Neelam Sinha

    Recent advances in neuroimaging have enabled studies in functional connectivity (FC) of human brain, alongside investigation of the neuronal basis of cognition. One important FC study is the representation of vision in human brain. The release of publicly available dataset BOLD5000 has made it possible to study the brain dynamics during visual tasks in great

  41. Yang Zhou, Ruixuan Zhu

    We prove the existence and regularity of convex solutions to the first initial-boundary value problem for the parabolic Monge-Amp\`ere equationn $$ \left\{\begin{eqnarray} &&-u_t+\det D^2u= \psi(x,t) \quad\quad\ \text{ in } Q_T,\newline &&u=\phi\quad\text{ on }\partial_pQ_T, \end{eqnarray}\right. $$ where $\psi,\phi$ are given functions, $Q_T=\Omega\times(0,

  42. Lu Zhang, Xue-Yang Song

    The moir\'e system provides a tunable platform for exploring exotic phases of materials. This article shows the possible realization of a non-Abelian state characterized by the Moore-Read wavefunction in a half-filled moir\'e Chern band, exemplified by twisted $\rm MoTe_2$. This is achieved by introducing short-range repulsive three-body interaction. Exact d

  43. Matthew Zurek, Yudong Chen

    We study the sample complexity of learning an $\varepsilon$-optimal policy in an average-reward Markov decision process (MDP) under a generative model. For weakly communicating MDPs, we establish the complexity bound $\widetilde{O}(SA\frac{H}{\varepsilon^2} )$, where $H$ is the span of the bias function of the optimal policy and $SA$ is the cardinality of th

  44. Shu-Min Wu, Xiao-Wei Teng, Xiao-Li Huang, Jianbo Lu

    Complex quantum information tasks in a gravitational background require multipartite entanglement for effective processing. Therefore, it is necessary to investigate the properties of multipartite entanglement in a relativistic setting. In this paper, we study genuine N-partite entanglement of massless Dirac fields in the Schwarzschild-de Sitter (SdS) spacet

  45. Meghna Menon, Devika Kamath, Maksym Mohorian, Hans Van Winckel

    Post-asymptotic giant branch stars (post-AGB) in binary systems, with typical orbital periods between ~100 to ~1000 days, result from a poorly understood interaction that terminates their precursory AGB phase. The majority of these binaries display a photospheric anomaly called 'chemical depletion', thought to arise from an interaction between the circumbina

  46. Taiga Adachi, Keiichiro Nomoto, Ryota Shii

    Let $f$ be a normalized newform of weight 2 on $\Gamma_0(N)$ whose coefficients lie in $\mathbb{Q}$ and let $\chi_M$ be a primitive quadratic Dirichlet character with conductor $M$. In this paper, under mild assumptions on $M$, we give a sharp lower bound of the 2-adic valuation of the algebraic central $L$-value $L(f, \chi_M, 1)$ and evaluate the 2-adic val

  47. Qinghua Zhao, Jiaang Li, Lei Li, Zenghui Zhou

    Existing works have studied the impacts of the order of words within natural text. They usually analyze it by destroying the original order of words to create a scrambled sequence, and then comparing the models' performance between the original and scrambled sequences. The experimental results demonstrate marginal drops. Considering this findings, different

  48. Minsu Kim, Jinwoo Hwang, Guseul Heo, Seiyeon Cho

    Learned indexes use machine learning models to learn the mappings between keys and their corresponding positions in key-value indexes. These indexes use the mapping information as training data. Learned indexes require frequent retrainings of their models to incorporate the changes introduced by update queries. To efficiently retrain the models, existing lea

  49. Pengyu Lai, Jing Wang, Rui Wang, Dewu Yang

    Predicting and understanding the chaotic dynamics in complex systems is essential in various applications. However, conventional approaches, whether full-scale simulations or small-scale omissions, fail to offer a comprehensive solution. This instigates exploration into whether modeling or omitting small-scale dynamics could benefit from the well-captured la

  50. Feng Shao, Dongyi Wei, Zhifei Zhang

    Motivated by recent breakthrough on smooth imploding solutions of compressible Euler, we construct self-similar smooth imploding solutions of isentropic relativistic Euler equations with isothermal equation of state $p=\frac1\ell\varrho$ for \textit{all} $\ell>1$ in physical space dimension $d=2,3$ and for $\ell>1$ close to 1 in higher dimensions. This work

  51. Jiaxu Zhang, Xin Chen, Gang Yu, Zhigang Tu

    Stylized motion breathes life into characters. However, the fixed skeleton structure and style representation hinder existing data-driven motion synthesis methods from generating stylized motion for various characters. In this work, we propose a generative motion stylization pipeline, named MotionS, for synthesizing diverse and stylized motion on cross-struc

  52. Siyu Xu, Yunke Wang, Daochang Liu, Bo Du

    Recent advancements in generative AI have suggested that by taking visual prompts, GPT-4V can demonstrate significant proficiency in visual recognition tasks. Despite its impressive capabilities, the financial cost associated with GPT-4V's inference presents a substantial barrier to its wide use. To address this challenge, we propose a budget-friendly collag

  53. Pablo Rocha

    Let $\mathbb{H}^{n}$ be the Heisenberg group. For $0 \leq \alpha < Q=2n+2$ and $N \in \mathbb{N}$ we consider exponent functions $p(\cdot) : \mathbb{H}^{n} \to (0, +\infty)$, which satisfies H\"older conditions, such that $\frac{Q}{Q+N} < p_{-} \leq p(\cdot) \leq p_{+} < \frac{Q}{\alpha}$. In this article we prove the $H^{p(\cdot)}(\mathbb{H}^{n}) \to L^{q(\

  54. Abdul Q. Batin, Suranjana Ghosh, Prasanta K. Panigrahi, Utpal Roy

    We report the close form expressions of the photon number statistics for a generalized coherent state and a generalized photon-added coherent state, which are shown to be crucial for proposing a variety of quantum scissor operations. The analytically obtained distributions are also capable of predicting the precise laser intensity windows for realizing a var

  55. Bosai Lyu, Jiajun Chen, Sen Wang, Shuo Lou

    Van der Waals encapsulation of two-dimensional materials within hexagonal boron nitride (h-BN) stacks has proven to be a promising way to create ultrahigh-performance electronic devices. However, contemporary approaches for achieving van der Waals encapsulation, which involve artificial layer stacking using mechanical transfer techniques, are difficult to co

  56. Ziru Niu, Hai Dong, A. K. Qin

    Personalized Federated Learning (PFL) is widely employed in IoT applications to handle high-volume, non-iid client data while ensuring data privacy. However, heterogeneous edge devices owned by clients may impose varying degrees of resource constraints, causing computation and communication bottlenecks for PFL. Federated Dropout has emerged as a popular stra

  57. Chaolei Tan, Jianhuang Lai, Wei-Shi Zheng, Jian-Fang Hu

    Video Paragraph Grounding (VPG) is an emerging task in video-language understanding, which aims at localizing multiple sentences with semantic relations and temporal order from an untrimmed video. However, existing VPG approaches are heavily reliant on a considerable number of temporal labels that are laborious and time-consuming to acquire. In this work, we

  58. Kanase Pankaj Popatrao

    There has been a significant effort in recent years to generalize the traditional concept of iterated function systems (IFS).In this article, we proposed Suzuki contraction in hyperspace and finding out the fixed point for Hutchinson mapping, which is called a deterministic fractal. The deterministic fractal for such a Suzuki contraction mapping is shown to

  59. Weiyao Wang, Yutian Lei, Shiyu Jin, Gregory D. Hager

    In this work, we introduce the Virtual In-Hand Eye Transformer (VIHE), a novel method designed to enhance 3D manipulation capabilities through action-aware view rendering. VIHE autoregressively refines actions in multiple stages by conditioning on rendered views posed from action predictions in the earlier stages. These virtual in-hand views provide a strong

  60. Teppei Suzuki

    In this work, we present Fed3DGS, a scalable 3D reconstruction framework based on 3D Gaussian splatting (3DGS) with federated learning. Existing city-scale reconstruction methods typically adopt a centralized approach, which gathers all data in a central server and reconstructs scenes. The approach hampers scalability because it places a heavy load on the se

  61. Yiwei Li, Zihao Wu, Huaqin Zhao, Tianze Yang

    To tackle the "reality gap" encountered in Sim-to-Real transfer, this study proposes a diffusion-based framework that minimizes inconsistencies in grasping actions between the simulation settings and realistic environments. The process begins by training an adversarial supervision layout-to-image diffusion model(ALDM). Then, leverage the ALDM approach to enh

  62. Vivek Singh, V. B. Tiwari, A. Chaudhary, S. Sarkar

    This study presents investigations on pulsed loading of a magneto-optical trap (MOT) on an atom chip in an UHV environment. Using three parallel resistively heated Rb-metal dispensers activated by pulsed current supply, approximately 3.0 $\times$ $10^{7}$ cold $^{87}Rb$ atoms were loaded into the MOT. A current pulse of $\sim$ 24 A with duration of $\sim$ 10

  63. Logan A. Burnett, Matthew P. Clay, Yogesh K. Vohra, Cheng-Chien Chen

    Using density functional theory (DFT) and linear response approaches, we compute the on-site Hubbard interaction $U$ of elemental Terbium (Tb) metal in the pressure range $\sim 0-65$ GPa. The resulting first-principles $U$ values with experimental crystal structures enable us to examine the magnetic properties of Tb using a self-consistent DFT+U method. The

  64. Huy Nghiem, Hal Daumé

    The widespread use of social media necessitates reliable and efficient detection of offensive content to mitigate harmful effects. Although sophisticated models perform well on individual datasets, they often fail to generalize due to varying definitions and labeling of "offensive content." In this paper, we introduce HateCOT, an English dataset with over 52

  65. Xin-Wei Yi, Ying Meng, Jia-Wen Li, Zheng-Wei Liao

    La$_{3}$Ni$_{2}$O$_{7}$ has garnered widespread interest recently due to its high-temperature superconductivity under pressure, accompanied by charge density wave (CDW) ordering and metal-insulator (MI) transitions in the phase diagram. Here, we reveal with comprehensive calculations that La$_{3}$Ni$_{2}$O$_{7}$ possesses an antiferromagnetic ground state un

  66. Ning Ning

    Expander graphs are fundamental in both computer science and mathematics, with a wide array of applications. With quantum technology reshaping our world, quantum expanders have emerged, finding numerous uses in quantum information theory, quantum complexity, and noncommutative pseudorandomness. The classical expander mixing lemma plays a central role in grap

  67. Hongrui Cai, Yuting Xiao, Xuan Wang, Jiafei Li

    We introduce a novel approach to creating ultra-realistic head avatars and rendering them in real-time (>30fps at $2048 \times 1334$ resolution). First, we propose a hybrid explicit representation that combines the advantages of two primitive-based efficient rendering techniques. UV-mapped 3D mesh is utilized to capture sharp and rich textures on smooth surf

  68. Muye Li, Shun Zhang, Yao Ge, Zan Li

    Integrated sensing and communication (ISAC) has become a promising technology for future communication system. In this paper, we consider a millimeter wave system over high mobility scenario, and propose a novel simultaneous transmission and reflection reconfigurable intelligent surface (STAR-RIS) aided ISAC scheme. To improve the communication service of th

  69. Haolan Chen, Jinhua Hao, Kai Zhao, Kun Yuan

    The objective of image super-resolution is to generate clean and high-resolution images from degraded versions. Recent advancements in diffusion modeling have led to the emergence of various image super-resolution techniques that leverage pretrained text-to-image (T2I) models. Nevertheless, due to the prevalent severe degradation in low-resolution images and

  70. Jiahe Wang, Jiale Huang, Bingzhao Cai, Yifan Cao

    Conventional approaches to facial expression recognition primarily focus on the classification of six basic facial expressions. Nevertheless, real-world situations present a wider range of complex compound expressions that consist of combinations of these basics ones due to limited availability of comprehensive training datasets. The 6th Workshop and Competi

  71. Hang Gao, Jiaguo Yuan, Jiangmeng Li, Peng Qiao

    Graph Neural Networks (GNNs) have garnered widespread attention for their potential to address the challenges posed by graph representation learning, which face complex graph-structured data across various domains. However, due to the inherent complexity and interconnectedness of graphs, accurately annotating graph data for training GNNs is extremely challen

  72. Li Tuobang

    As the most fundamental problem in statistics, robust location estimation has many prominent solutions, such as the trimmed mean, Winsorized mean, Hodges Lehmann estimator, Huber M estimator, and median of means. Recent studies suggest that their maximum biases concerning the mean can be quite different, but the underlying mechanisms largely remain unclear.

  73. Linyu Tang, Lei Zhang

    Numerous studies have demonstrated the susceptibility of deep neural networks (DNNs) to subtle adversarial perturbations, prompting the development of many advanced adversarial defense methods aimed at mitigating adversarial attacks. Current defense strategies usually train DNNs for a specific adversarial attack method and can achieve good robustness in defe

  74. Zhiyang Guo, Wengang Zhou, Li Li, Min Wang

    3D Gaussian Splatting (3DGS) has become an emerging tool for dynamic scene reconstruction. However, existing methods focus mainly on extending static 3DGS into a time-variant representation, while overlooking the rich motion information carried by 2D observations, thus suffering from performance degradation and model redundancy. To address the above problem,

  75. Clint Morris, Michael Jurado, Jason Zutty

    In the realm of machine learning, traditional model development and automated approaches like AutoML typically rely on layers of abstraction, such as tree-based or Cartesian genetic programming. Our study introduces "Guided Evolution" (GE), a novel framework that diverges from these methods by utilizing Large Language Models (LLMs) to directly modify code. G

  76. Bo Jiang, Jian Du, Sagar Sharma, Qiang Yan

    Differential Privacy (DP) mechanisms usually {force} reduction in data utility by producing "out-of-bound" noisy results for a tight privacy budget. We introduce the Budget Recycling Differential Privacy (BR-DP) framework, designed to provide soft-bounded noisy outputs for a broad range of existing DP mechanisms. By "soft-bounded," we refer to the mechanism'

  77. Linlin Huang, Yuanyuan Wang, He-Xu Zhang, Shinya Matsuzaki

    We argue that the axionic domain-wall with a QCD bias may be incompatible with the NANOGrav 15-year data on a stochastic gravitational wave (GW) background, when the domain wall network collapses in the hot-QCD induced local CP-odd domain. This is due to the drastic suppression of the QCD bias set by the QCD topological susceptibility in the presence of the

  78. Dipayan Banerjee, Alan Erera, Alejandro Toriello

    We study the finite-horizon continuous-time dynamic yield management problem with stationary arrival rates and two customer types. We consider a class of linear threshold policies proposed by Hodge (2008), in which each less-profitable customer is accepted if and only if the remaining inventory exceeds a threshold that linearly decreases over the horizon. We

  79. Guohang Zhuang, Yue Hu, Tianxing Yan, JiaZhan Gao

    Currently, most food recognition relies on deep learning for category classification. However, these approaches struggle to effectively distinguish between visually similar food samples, highlighting the pressing need to address fine-grained issues in food recognition. To mitigate these challenges, we propose the adoption of a Gaussian and causal-attention m

  80. Masaki Tsukamoto

    The main purpose of this paper is to propose an ergodic theoretic approach to the study of entire holomorphic curves. Brody curves are one-Lipschitz holomorphic maps from the complex plane to the complex projective space. They naturally form a dynamical system, and "random Brody curves" in the title refers to invariant probability measures on it. We study th

  81. Xu Jing, Cheng Qian, Chen-Xun Weng, Bing-Hong Li

    Quantum communication networks are crucial for both secure communication and cryptographic networked tasks. Building quantum communication networks in a scalable and cost-effective way is essential for their widespread adoption, among which a stable and miniaturized high-quality quantum light source is a key component. Here, we establish a complete polarizat

  82. Weiwei Zhou, Jiada Lu, Chenkun Ling, Weifeng Wang

    Human emotion recognition holds a pivotal role in facilitating seamless human-computer interaction. This paper delineates our methodology in tackling the Valence-Arousal (VA) Estimation Challenge, Expression (Expr) Classification Challenge, and Action Unit (AU) Detection Challenge within the ambit of the 6th Workshop and Competition on Affective Behavior Ana

  83. Jinpeng Li, Zekai Zhang, Quan Tu, Xin Cheng

    Large Language Models (LLMs) demonstrate superior performance in generative scenarios and have attracted widespread attention. Among them, stylized dialogue generation is essential in the context of LLMs for building intelligent and engaging dialogue agent. However the ability of LLMs is data-driven and limited by data bias, leading to poor performance on sp

  84. Abel Dasylva, Arthur Goussanou, Christian-Olivier Nambeu

    The capture-recapture method can be applied to measure the coverage of administrative and big data sources, in official statistics. In its basic form, it involves the linkage of two sources while assuming a perfect linkage and other standard assumptions. In practice, linkage errors arise and are a potential source of bias, where the linkage is based on quasi

  85. Zhengtang Tan, Shouchuan Zhang

    All quasi-affine connected Generalized Dynkin Diagrams with rank $> 5$ are found. All quasi-affine Nichols (Lie braided) algebras with rank $> 5$ are also found.

  86. Chenyi Li, Ziyu Wang, Wanyi He, Yuxuan Wu

    The convergence rate of various first-order optimization algorithms is a pivotal concern within the numerical optimization community, as it directly reflects the efficiency of these algorithms across different optimization problems. Our goal is making a significant step forward in the formal mathematical representation of optimization techniques using the Le

  87. Nabarun Goswami, Yusuke Mukuta, Tatsuya Harada

    The success of models operating on tokenized data has heightened the need for effective tokenization methods, particularly in vision and auditory tasks where inputs are naturally continuous. A common solution is to employ Vector Quantization (VQ) within VQ Variational Autoencoders (VQVAEs), transforming inputs into discrete tokens by clustering embeddings in

  88. Weijun Fang, Jingke Xu, Ruiqi Zhu

    The deep holes of a linear code are the vectors that achieve the maximum error distance (covering radius) to the code. {Determining the covering radius and deep holes of linear codes is a fundamental problem in coding theory. In this paper, we investigate the problem of deep holes of twisted Reed-Solomon codes.} The covering radius and a standard class of de

  89. Yifan Wang, Yafei Liu, Chufan Shi, Haoling Li

    Instruction tuning effectively optimizes Large Language Models (LLMs) for downstream tasks. Due to the changing environment in real-life applications, LLMs necessitate continual task-specific adaptation without catastrophic forgetting. Considering the heavy computational cost, replay-based Continual Learning (CL) methods are the simplest and most widely used

  90. Kuntai Du, Yihua Cheng, Peder Olsen, Shadi Noghabi

    With the increasing deployment of earth observation satellite constellations, the downlink (satellite-to-ground) capacity often limits the freshness, quality, and coverage of the imagery data available to applications on the ground. To overcome the downlink limitation, we present Earth+, a new satellite imagery compression system that, instead of compressing

  91. Ajai Choudhry, Arman Shamsi Zargar

    Since 1772, when Euler first described two methods of obtaining two pairs of biquadrates with equal sums, several methods of solving the diophantine equation $x^4+y^4=z^4+w^4$ have been published. All these methods yield parametric solutions in terms of homogeneous bivariate polynomials of odd degrees. In this paper we describe a method that yields three par

  92. Farhad Farokhi, Sejeong Kim

    Gentle quantum leakage is proposed as a measure of information leakage to arbitrary eavesdroppers that aim to avoid detection. Gentle (also sometimes referred to as weak or non-demolition) measurements are used to encode the desire of the eavesdropper to evade detection. The gentle quantum leakage meets important axioms proposed for measures of information l

  93. Wenjie Zhang, Yuxiang Wan, Zhong Zhuang, Ju Sun

    For nonlinear inverse problems that are prevalent in imaging science, symmetries in the forward model are common. When data-driven deep learning approaches are used to solve such problems, these intrinsic symmetries can cause substantial learning difficulties. In this paper, we explain how such difficulties arise and, more importantly, how to overcome them b

  94. Hanxi Wan, Pei Li, Arpan Kusari

    With the advent of universal function approximators in the domain of reinforcement learning, the number of practical applications leveraging deep reinforcement learning (DRL) has exploded. Decision-making in autonomous vehicles (AVs) has emerged as a chief application among them, taking the sensor data or the higher-order kinematic variables as the input and

  95. Yusuke Kimura, Tomotaka Kuwahara

    This paper delves into a fundamental aspect of quantum statistical mechanics -- the absence of thermal phase transitions in one-dimensional (1D) systems. Originating from Ising's analysis of the 1D spin chain, this concept has been pivotal in understanding 1D quantum phases, especially those with finite-range interactions as extended by Araki. In this work,

  96. Jiaxin Guo, Hao Yang, Zongyao Li, Daimeng Wei

    This paper presents a study on strategies to enhance the translation capabilities of large language models (LLMs) in the context of machine translation (MT) tasks. The paper proposes a novel paradigm consisting of three stages: Secondary Pre-training using Extensive Monolingual Data, Continual Pre-training with Interlinear Text Format Documents, and Leveragi

  97. Jiancheng Zhao, Jiaqi Yue, Chunhui Zhao

    Zero-shot fault diagnosis (ZSFD) is capable of identifying unseen faults via predicting fault attributes labeled by human experts. We first recognize the demand of ZSFD to deal with continuous changes in industrial processes, i.e., the model's ability to adapt to new fault categories and attributes while avoiding forgetting the diagnosis ability learned prev

  98. Sebin Oh, Sang-ri Yi, Ziqi Wang

    This study introduces the long-range Ising model from statistical mechanics to the Performance-Based Earthquake Engineering (PBEE) framework for regional seismic damage analysis. The application of the PBEE framework at a regional scale involves estimating the damage states of numerous structures, typically performed using fragility function-based stochastic

  99. Nicolò Dalmasso, Antonello Calabrò, Nicha Leethochawalit, Benedetta Vulcani

    We present an analysis of the galaxy merger rate in the redshift range $4.0<z<9.0$ (i.e. about 1.5 to 0.5 Gyr after the Big Bang) based on visually identified galaxy mergers from morphological parameter analysis. Our dataset is based on high-resolution NIRCam JWST data (a combination of F150W and F200W broad-band filters) in the low-to-moderate magnification

  100. Tingyang Zhang, Qingzhe Gao, Weiyu Li, Libin Liu

    Animatable 3D reconstruction has significant applications across various fields, primarily relying on artists' handcraft creation. Recently, some studies have successfully constructed animatable 3D models from monocular videos. However, these approaches require sufficient view coverage of the object within the input video and typically necessitate significan