Skip to content

Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?

Alberto Cetoli

cs.CLarXiv:2608.29921

Abstract

The output of a Language Model can be tampered with while the model is writing it. A simple test can thus be constructed by evaluating the model's perception of this external perturbation. In this spirit, a simple benchmark is built in which a single word is consistently substituted with another in the generation process. We call this method Sleight of Word. Two distinct axes are measured: metrics that relate to the model's surprise, as well as an evaluation of the textual reaction for 19 different open-weight language models.

Create a lesson