What survives honest evaluation? Leakage-safe, search-aware assessment of LLM-driven trading strategy discovery
Eray Gençay
Abstract
Large language models (LLMs) are increasingly used to discover trading strategies, and much of the resulting literature shares a methodological weakness: many candidate strategies are generated, the best is reported, and neither look-ahead bias nor the intensity of the search behind the reported result is corrected for. We present a strategy-discovery system that makes both corrections structural rather than procedural. First, the agent can only act through registry-validated tools whose feature space excludes look-ahead by construction; we show that this guardrail is not redundant with statistical correction: a deliberately leaky oracle posting a Sharpe ratio of 35 survives Deflated Sharpe and probability-of-backtest-overfitting testing completely. Second, the system records every strategy evaluation its search performs and deflates all reported performance by that trial count, tracing how the best in-sample Sharpe ratio climbs with each trial while the deflation threshold, driven by the agent's own search, climbs faster. Across a 453-stock point-in-time US equity universe and a 39-ETF multi-asset universe with realistic transaction, impact, and borrow costs, honest evaluation certifies passive benchmarks (out-of-sample confidence intervals excluding zero), rejects every LLM-discovered strategy (across two frontier models, search budgets up to one hundred candidates, and five repeated runs), catching selection luck, predicted rank degradation, and out-of-sample collapse through complementary instruments, and evaluates a human trader's production rule system under identical instruments. The framework formalizes why pre-registered hypotheses earn lower evidential bars than brute search, and quantifies the sample sizes that credible certification of moderate edges actually requires.
Create a lesson
Related papers
Modeling Trade Durations under Temporal Granularity Effects in Forex Markets
Vladimír Holý
Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations
Marcus Gawronsky, Chun-Sung Huang
Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations
Marcus Gawronsky, Chun-Sung Huang
Deep Hedging Under Realistic Market Frictions: A Regime-Conditional Empirical Study of Dynamic Option Hedging on Bitcoin Options
Sheryan Kumar
Lead-Lag Relationships in Financial Markets: A Comparison of Multiple Clustering Algorithms
Ruichen Deng, Yichi Zhang
Equity Strategy Backtesting: Luck or Edge? The MinervaScore as a Statistical Robustness Grade
Maria Laura Santoni, Vincent Jouanne, Matthew L. Scullin