Evaluating for the long term: Learnings from industry
Leif Sigerson, Tom Cunningham, Winston Chou, Sana Pandey, Jonathan Stray, Lo-Hua Yuan, Eytan Bakshy, Timothy Chan, Molly Davies, Maria Dimakopoulou, Simon Ejdemyr, Kenneth Hung, Nathan Kallus, Madhav Kumar, Thu Le, A. Demetri Pananos, Lee Richardson, Brennan Schaffner, Rose Tan, Martin Tingley, Nadia Tomova, Panagiotis Toulis, Wenjing Zheng, Zander Arnao, Dean Eckles
Abstract
Online platforms prioritize long-term business outcomes, yet typical experiments are far too short to measure these outcomes directly. Our goal in this paper is to collect and share industry knowledge on how to make decisions from short-term experiments that are better aligned with long-term outcomes. Based on a daylong workshop with 26 experts from 15 online platforms and 4 universities, we formulate a series of propositions that reflect current industry knowledge. Participants largely agreed that reversals of sign from short-run to long-run treatment effects are rare, with reversals concentrating in specific cases such as treatments involving content quality signals, hyper-monetization, and pricing. Although the magnitude of treatment effects can shift over time, a "univariate autosurrogate", corresponding to the short-run treatment effect on the long-run metric of interest, is often hard to beat. A recurring theme was the importance of surrogates that are not only (or even primarily) unbiased for true long-run outcomes, but that improve decision-making. Thus, participants generally agreed that simple, interpretable surrogates were generally preferable to elaborate but hard-to-explain surrogate indices. Participants also agreed that, due to concerns about confounding and transportability, experimentally-learned surrogates are generally preferable to observationally-learned surrogates. However, the drawback is that learning good surrogates from experiments typically requires a large, representative portfolio of long-run experiments that few platforms possess. We conclude that there is no substitute for a well-run long-term experiment, whether for learning surrogates or validating them, and we highlight open challenges including evolving treatments, persistent treatments not fully mediated by short-term proxies, and mismatch between experimental samples and the target population.
Create a lesson
Related papers
Characterising mortality dynamics across countries and time using a multi-stage clustering approach
Pedro Menezes de Araújo, Ugofilippo Basellini, Thomas Brendan Murphy et al.
Combining Weather Forecast Aggregation and State-Space Models for Adaptive Probabilistic Electricity Load Forecasting
Joseph de Vilmarest, Jonathan Dumas, Jean Thorey
Anthropogenic Forcing, Climate Change, and the Shape of Warming: Statistical Inference for Distributional Cointegration
Won-Ki Seo, Kyungsik Nam
Observational constraints on net radiative forcing confirm aviation contrail warming
Aaron Sonabend-W, Scott Geraedts, Nita Goyal et al.
Nationally Consistent, Locally Incomplete: A Bayesian Remote-Sensing Audit of Rooftop Photovoltaic Registries
Gabriel Kasmi, Yves-Marie Saint-Drenan, Laurent Dubus et al.
Temporal Seam Score for Assessing Continuity at Known Transitions in Time Series
Hongxiao Jin