Beyond Ranking Accuracy: Evaluating LLM-Cited Feature Rationales for Next Basket Repurchase Recommendation
Yanan Cao, Anay Dombe, Murali Mohana Krishna Dandu, Shreeranjani Srirangamsridharan, Sinduja Subramaniam, Yogananth Mahalingam, Evren Korpeoglu, Kannan Achan
Abstract
Next-basket repurchase recommendation is commonly formulated as a ranking task: given a customer's purchase history, the system ranks previously purchased items that may be needed again. In production settings, however, ranking accuracy is only one component of recommendation quality. Customers may also benefit from concise evidence about why an item is recommended now. Large language models (LLMs) offer a potential way to surface such evidence through feature-based, human-readable rationales grounded in interpretable behavioral signals. We construct repurchase features spanning cadence, frequency, recency, user behavior, and item popularity, and evaluate LLMs on two public grocery datasets and one proprietary retail dataset. We investigate (1) whether off-the-shelf LLMs can use these features as next-basket scorers relative to heuristic and supervised rankers, and (2) whether LLM-cited features carry outcome-grounded ranking signal. For the latter, we compare LLM-cited features with model-specific attribution methods under a cross-model feature-masking protocol that measures ranking degradation after masking selected features. Our results show that LLM scores are not competitive with supervised rankers, suggesting that off-the-shelf LLMs should not be used as standalone repurchase recommenders. However, changes in prompt and evidence representation can improve outcome-grounded feature-masking results in some settings even when ranking performance does not improve; the effect is dataset-dependent and does not consistently match attribution baselines. These findings suggest a practical role for LLMs as validated explanation components rather than primary rankers, with rationale quality evaluated separately from ranking accuracy.
Create a lesson
Related papers
Closed Forms and Synthetic Twins: Predicting Approximate Nearest Neighbor Recall from Embedding Statistics
Shmuel Herman
MUSES: A Benchmark for Prospective Intellectual-Roots Retrieval
Rohan Pandey, Sunjae Kwon, Hong Yu
Two-Sided State-Space Models for Sequential Recommendation with Non-Random Multimodal Review Feedback
Ziwen Pan, Zihan Liang, Ruoxuan Xiong
MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval
Seokwon Song, Sohyeon Kim, Gunhee Kim
Learning from What You Retrieve: Online RL Fine-Tuning for Semantic Retrieval
Shaowei Wei, Chong Huang, Songtao Fang et al.
Generative Retrieval for E-commerce: Jointly Learning Embedding and Codebook with Same Product Cluster
Songtao Fang, Zihao Xu, Shaowei Wei et al.