Skip to content

What Does a Benford Test Actually Test? Marginal Conformity, Sampling Structure, and Forensic Inference

Arthur Charpentier

stat.MEarXiv:2609.18424

Abstract

Benford's law specifies a marginal distribution for significant digits, whereas the usual first-digit Pearson p-value is calibrated under an independent multinomial sampling model. We separate these statements with four constructions that share the same one-time or pooled Benford target but have different joint structures. Under a wrapped-Gaussian circular Markov construction, the nominal 5% Pearson test rejects 21.6% of samples at a fixed persistence level even though every one-time marginal is exactly Benford; a sequence-level second-order moment-matching calibration reduces the rejection rate to 4.8%. The same strategy performs well in the randomized-rotation and random-composition designs. The effect persists across changes in sample size and sequence length, and a grid approximation to a continuous log-significand Cramér-von Mises discrepancy shows the same design dependence. In the wrapped-Gaussian design examined here, corrected procedures retain rejection probabilities that rise with the size of a smooth marginal departure. Finally, we distinguish calibration from an information boundary: significand-only methods cannot detect changes that leave the complete significand process unchanged, but they readily detect perturbations that alter it. Benford conformity is a marginal statement; a Benford p-value is valid only relative to a specified statistic, sampling law, and calibration procedure; an integrity claim requires substantive competing models.

Create a lesson