Big data, differential privacy, and national statistical organisations
James Bailie
Abstract
Differential privacy (DP) has emerged in the computer science literature as a measure of the impact on an individual's privacy resulting from the publication of a statistical output such as a frequency table. This paper provides an introduction to DP for official statisticians and discuss its relevance, benefits, and challenges from a National Statistical Organisation (NSO) perspective. We motivate our study by examining how privacy is evolving in the era of big data and how this might prompt a shift from traditional statistical disclosure techniques used in official statistics--which are generally applied on a cell-by-cell or table-by-table basis--to formal privacy methods, like DP, which are applied from a perspective encompassing the totality of the outputs generated from a given dataset. We identify an important interplay between DP's holistic privacy risk measure and the difficulty for NSOs in implementing DP, showing that DP's major advantage is also DP's major challenge. This paper provides new work addressing two key DP research areas for NSOs: DP's application to survey data and its incorporation within the Five Safes framework.
Create a lesson
Related papers
Learning CNN Filters via Generalized Stein's Method
Guang Yang, Wei Shi, Yuan Cao et al.
Did Mary Shelley Write Frankenstein? A Stylometric Analysis
Lee Suddaby, Gordon J Ross
Sequential Certification of Threshold Decisions in Rare-Event Risk Prediction
Hui-Mean Foo, Yuan-chin Ivan Chang
Neural Networks Learning the Radon--Nikodym Derivative: Empirical Option Pricing in Incomplete Markets
Ziyuan Zhang, Kiseop Lee
Random Forest-Informed Cellular Automaton for Large-Scale Wildfire Spread Modelling
Siyu Chen, Esha Saha, Hao Wang
A Tool for Reconstructing Transit Vehicle Trajectories: A Case Study at IndyGo
Ben O'Brien, Lewis J. Lehe