TabJoinBench: A Benchmark for Joinable Table Discovery
Sandipan De, Jin Wang, Vivek Gupta
Abstract
Join discovery aims to identify tables from large data repositories that can augment a query table with complementary information, enabling downstream tasks such as data exploration, feature engineering, and business intelligence. Although numerous join discovery methods have been proposed, existing studies rely on method-specific benchmark construction, making reproducible and fair comparison difficult. We present TabJoinBench, a benchmark for evaluating join discovery methods across semantic, relational, and hybrid data lake scenarios. TabJoinBench constructs query-candidate pairs using source-specific validation strategies, systematically introduces structural, representation, and semantic changes through composable perturbations while preserving reliable ground truth. We evaluate representative join discovery methods spanning set-based, feature-based, and learned approaches, together with general-purpose language-model embedding baselines, and publicly release the processed datasets, ground-truth annotations, and generation pipeline to facilitate reproducible evaluation and future research.
Create a lesson
Related papers
Prune First, Decide Fast: Scalable Semantic Query Processing with JEVDB
Zhengle Wang, Hanxu Yan, Fuheng Zhao et al.
DIADA: Automatic Data Composition in Data Lakes
Marc Maynou, Albert Martin, Sergi Nadal et al.
Bridging the Omics Divide: A Modular Relational Approach to Multi-Layer Biological Data Management
Alessandro Balestrucci, Donald Friggieri, Andrea Gariboldi et al.
STEER: Reducing Inference Cost in Relational Foundation Models through Semantically Informed Sampling
Abdalla Mohamed, Ashraf Aboulnaga
HakiCC: LLM-Driven Multi-Agent Design and Optimization of Concurrency Control Protocols
Farzad Habibi, Juncheng Fang, Faisal Nawab
On Enforcing Database Constraints Using MS VBA Event Driven Procedures and SQL Server
Diana Christina Mancas