Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
Junghyun Min, Huseyin Uzunalioglu, Mohamed Trabelsi
Abstract
Recent breakthroughs in LLM-based systems and their abilities in problem solving and coding have allowed progress in the AI for Science paradigm, potentially replacing human roles in machine learning (ML) research. However, while several frameworks of fully autonomous end-to-end ML research have been proposed, successful implementations of them are often limited to problems with narrow search spaces, like language modeling or biomedical ML benchmarks. In this paper, we explore how autonomous research can be adapted to solve open-ended, industry-grade ML problems, by considering a case study: telecom ticket retrieval, an open-ended task with degrees of freedom in representation, architecture, and training data generation. We discover that autonomous research for open-ended problems with commercial and open-source agents shows both promise and limitations: while autonomous research can excel in narrow hyperparameter optimization, it lacks human-like intuition and creativity and requires operational overhead. Even with minimal human supervision, autonomous research can reach 90\% of state-of-the-art performance (0.34 vs. 0.38 Recall@1) in a much shorter time period (10 weeks vs. 10 months of human work) at a modest cost (up to \$200 per Cursor campaign). Our empirical evidence recommends that human researchers and autonomous research frameworks work together for best results in ML research.
Create a lesson
Related papers
Rethinking Heterogeneous System Disaggregation for Subquadratic Attention
Arya Tschand, Yaosheng Fu, Vikram Sharma Mailthody et al.
A Hybrid LSTM-XGBoost Framework for Multi-Horizon Stock Return Prediction Across Diversified Equity Portfolios
Seif ElDein Mostafa, Yahia Ahmed, Farah Datwish et al.
CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
Jinting Wang, Chenxing Li, Dong Yu et al.
Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction
Baoyang Jiang, Fengchun Zhang, Leyuan Wang et al.
Anchoring Clinical Events in Time: UID-Preserving Multimodal Reconstruction and Source-Grounded Adjudication
Sayantan Kumar, Nicolas Grimaldi, Jack Cummins et al.
Diffusion Models and Concept Formation
Zekun Wang, Karthik Singaravadivelan, Christopher J. MacLellan