RENSA: Rich Environment Metadata to Navigate Shared and Distributed Endpoints for Automated Federated SPARQL Query Generation
Victor Eiti Yamamoto, Takeda Hideaki, Yamamoto Yasunori
Abstract
The number of knowledge graph databases has increased significantly with the proliferation of knowledge graph technologies. Knowledge graphs enable the dynamic integration of distributed data through federated SPARQL queries. However, constructing efficient queries in a federated environment is challenging due to the lack of detailed structural knowledge across decentralized datasets. While standards like VoID provide basic metadata, they often fail to capture the complex interlinks and authority distributions necessary for optimization. Consequently, current engines frequently rely on runtime ASK queries for source selection, increasing communication overhead. We propose RENSA, a federated SPARQL query generation framework that leverages an extension of SPARQL Builder Metadata (SBM). By integrating class and authority information, mapping subject and object usage to specific predicates, RENSA enables precise source selection and semantic constraint inference for query variables without runtime communication. The generated profiles represent less than 1\% of the original dataset triples in most cases, ensuring storage efficiency. Evaluation on the LargeRDFBench benchmark (13 datasets with >1B triples, 32 queries) shows that RENSA achieves source selection results comparable to state-of-the-art methods while eliminating ASK query overhead. Furthermore, we demonstrate that RENSA infers class and authority constraints for query variables, enabling the identification of data sources even across heterogeneous endpoints. These profiles additionally offer human-readable structural insights for semi-automated query generation.
Create a lesson
Related papers
Detecting DBMS Bugs by Constructing Equivalent Representations of Intermediate Query Results
Xiaoxu Niu, Gong Chen, Jinfu Chen et al.
ELASTIC: Trajectory-Based Synchronization of Event and Tracking Data in Soccer
Hyunsung Kim, Hoyoung Choi, Kunhee Lee et al.
Demystifying and Improving Lazy Promotion in Cache Eviction
Qinghan Chen, Muhammad Haekal Muhyidin Al-Araby, Ziyue Qiu et al.
Diachronic Hypergraphs for Orchestrated Multi-Agent Multimodal Memory Curation
Yichao Feng, Ran Zhang, Haoran Luo et al.
Engaging the scientific community in high-quality biocuration: a report on the International Society for Biocuration workshop, 'Maximizing community curation for the benefit of all'
Daniela Raciti, Susan L. M. Coort, Christian Grove et al.
No Silver Bullet: Boosting GaussDB Performance on the 30TB TPC-H Workload
Tim Zeyl, Jason Lam, Shu Lin et al.