From language-model stock rankings to testable economic rules: A computational audit
Shuai Wu, Xue Li, Zhijun Wang, Bolun Liu, Weilin Cai, Zihao Su, Ran Wang
Abstract
We test the stability, reproducibility and investment outcomes of language-model stock rankings. Four models and five numerical comparators share a portfolio engine over 72 monthly holding periods in the Shanghai Stock Exchange (SSE) 50, China Securities Index (CSI) 300 and CSI 500. Rankings use nine characteristics, and five repeated SSE 50 runs measure variation under identical inputs. Linear rules fitted to development-period model preferences are frozen before unseen-month, larger-pool and controlled-intervention tests. Their mean Spearman agreement with model rankings is 0.923-0.984 in the SSE 50 and 0.795-0.985 after transfer. Aggregate rank-change error falls relative to a zero-change prediction in twelve archived feature-group comparisons and eight matched single-feature comparisons, with Holm adjustments applied in separate nine-plus-three and six-plus-two families. Prediction of individual entries and exits remains weak (event Jaccard 0.000-0.125). Historical mean model compound annual growth rates range from 6.26% to 11.17%. At 10 basis points per side and six-month blocks, the twelve-comparison model-minus-rule return family and factor-controlled associations yield no adjusted finding. Three higher-cost, twelve-month-block comparisons favor a Terra rule within their twelve-test slices, with no adjusted finding across the full 144-test sensitivity grid. Two input-intervention batches totaling 5,184 responses supply paired intervention-return tests. The three-model batch has no bootstrap-adjusted finding at the primary block length; a Luna row-order effect appears under heteroskedasticity- and autocorrelation-consistent (HAC) adjustment within its three-test family but not in a pooled 24-test adjustment. Compact rules approximate aggregate rankings; return conclusions depend on comparison families, uncertainty methods and tie priorities.
Create a lesson
Related papers
RipplePLM: Structural and Property Decoupling for Protein Mutation Effect Generation
Liuzhenghao Lv, Yuyang Liu, Yuyang Gao et al.
CODesign: Consistency from Data to Trajectory in All-Atom Protein Binder Co-Design
Yuanle Mo, Bo Qiang, Haitao Lin et al.
A pyramidal ISRU lunar habitat design based on topological interlocking of sintered regolith blocks
Lukas Schnelle, Kai-Uwe Schröder, Alice C. Niemeyer et al.
Names without information: Attention allocation and rent transfer in a zero-fundamental token market
Dingding Cao, Han Wang, Yujing Zhong et al.
How Evaluation Choices Change the Measured Benefit of Cooperative Perception: Evidence from Three V2X Benchmarks
Pincan Zhao, Yili Tang, Xinrui Zhang
Trusting the Inverse: Reliability-Aware Mapping for Simulation-Based Microstructure Estimation in Diffusion MRI
Juan Luis Villarreal Haro, Ileana Jelescu, Jean-Philippe Thiran et al.