Who Wrote This? Turing, Total Variation, and the Mathematics of AI-Text Detection
Santiago Schnell
Abstract
What can a finished text reveal about the process that produced it? Drawing on Turing's imitation game and statistical decision theory, this article examines the limits of AI-text detection as an inference from a completed object to an unobserved history. For two known source distributions with equal prior probabilities, a standard identity gives the minimum average classification error as half their probability overlap. This overlap is one minus their total variation distance. Perfect detection therefore requires nonoverlapping distributions; useful discrimination does not. In practice, the problem is harder: human and machine writing form changing families of distributions, posterior probabilities depend on base rates, and ``AI-written'' becomes ambiguous when people and software contribute to the same text. Turing's interrogator can ask another question; a detector restricted to finished prose cannot. Distinguishability, source attribution, authentication, and compliance are different inferential tasks. The paper's own documented human-AI provenance shows what a production history can reveal that a binary label cannot.
Create a lesson
Related papers
The suppression thesis in Riemann's Habilitationsvortrag
Victor Tapia
Remembering Solomon Marcus
Florin Nichita
The Erdős--Sós Theorem
David R. Wood
Math for AI safety: an invitation for mathematicians
Lionel Levine
Immensely Questionable Tests
Siyona Agarwal, Julian Bernhoft, Karam Gill et al.
On the Reconstruction of SAS from Other Triangle Congruence Criteria
Roberto Volpe