Does the way we write a theory change the program an LLM builds from it? A prospective randomized study of renderer format in LLM theory-to-program translation
Andre Panossian
Abstract
A verbal theory does not run: translating it into an executable model requires choices about variables, interventions, and interactions. We tested whether the presentation of otherwise identical theoretical content systematically changes the programs produced by large language models. In a prospective preregistered randomized experiment, 16 independent assignment bits allocated 32 paired renderer slots between a structured intervention contract and connected prose. Two pinned LLM snapshots translated five anonymous theoretical accounts, yielding 320 preauthorized single-shot programs in a frozen sparse quadratic language. A deterministic evaluator measured atomic finite-difference responses (H1) and mixed interaction responses (H2). Two primary endpoints assessed cross-model matched-distance reduction and closed-set same-account identifiability, with exact randomization inference and 27 preregistered support criteria evaluated across signed-linear and magnitude-rank pipelines. Both H1 and H2 returned the registered verdict NOTSUPPORTED; only 19 of 108 criterion evaluations passed. Same-account identifiability remained near chance (AUC 0.469-0.523, against a registered 0.80 threshold). One H2 matched-distance endpoint moved and survived multiplicity correction in the signed-linear pipeline, but the corresponding magnitude-rank result missed the registered effect-size floor, so it did not satisfy the joint support rule. Thus renderer format did not produce the uniform, family-invariant, classifiable behavioral geometry predicted in advance. The result places a concrete boundary on specification-format effects in LLM theory-to-program translation and provides a fully auditable randomized design, datasets, and software for studying executable formalization.
Create a lesson
Related papers
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
Jeonghye Kim, Minseon Kim, Young Jin Kim et al.
Evaluating the Health of Open-Source Smart City Platforms
Rodrigo Bravo Simões, Fernando Brito e Abreu, Vasco Amaral
From Component Snapshots to Lifecycle Traces: Agent-Based Software Composition Analysis
Chaofan Li, Zhengduo Xue, Chengxiang Li et al.
A Study on the Impact of Natural Language Differences in Prompts on Automatic Code Generation Using LLMs
Haruka Tokumasu, Masanari Kondo, Alexander Serebrenik et al.
A Study of the Reliability of Agentic AI-Generated Programs
Ayesha Shafique, Barton P. MIller, Elisa R. Heymann
Relationally Guided Use Case Modeling with LLMs
Guangyu Wang, Bangqi Li, Ji Wu et al.