Designing a Multi-petabyte Database for LSST
Jacek Becla, Andrew Hanushevsky, Sergei Nikolaev, Ghaleb Abdulla, Alex Szalay, Maria Nieto-Santisteban, Ani Thakar, Jim Gray
Abstract
The 3.2 giga-pixel LSST camera will produce approximately half a petabyte of archive images every month. These data need to be reduced in under a minute to produce real-time transient alerts, and then added to the cumulative catalog for further analysis. The catalog is expected to grow about three hundred terabytes per year. The data volume, the real-time transient alerting requirements of the LSST, and its spatio-temporal aspects require innovative techniques to build an efficient data access system at reasonable cost. As currently envisioned, the system will rely on a database for catalogs and metadata. Several database systems are being evaluated to understand how they perform at these data rates, data volumes, and access patterns. This paper describes the LSST requirements, the challenges they impose, the data access philosophy, results to date from evaluating available database technologies against LSST requirements, and the proposed database architecture to meet the data challenges.
Create a lesson
Related papers
Compositional Online Learning for Semantic Data Processing Systems
Paweł Liskowski, Fuheng Zhao, Benjamin Han et al.
Incremental Delta-Shapley: A Standalone Runtime for Predicate Attribution on Sliding Windows
Pouya Khani, Ira Assent
Size Bounds for CQs Under Acyclic Constraints
Stefan Mengel, Andrei Romashchenko
IBLTs Measure Before They Decode: Self-Sizing Set Reconciliation from Pre-Peeling Counts
Min Wu, Ji Qi, Chengdui Luo et al.
VoS: Variate Ordering Strategies for Skyline Query Optimization
Abhinav Gorantla, Pratanu Mandal, K. Selçuk Candan et al.
Realistic Counterfactual Explanations via Denial Constraints
Avia Asael, Nave Frost, Amir Gilad et al.