Learning to structure data from user-generated thematic corpora
Elishay Avram, Oren Glickman, Elad Yom-Tov
Abstract
Thematic corpora, such as social media communities, contain unstructured text describing data that could be made structured. These include, for example, personal attributes, behaviors, and experiences mentioned in social media data. Extracting structured data is challenging as relevant attributes are often implicit, domain-dependent, and unknown in advance. We propose a fully automated, iterative framework for discovering and extracting domain-specific attribute schemas without a predefined ontology. Using large language models (LLMs), the framework induces candidate attributes, sequentially consolidates semantically overlapping attributes, and assigns a structural type. These enable creating an ontology and populating it with values from the corpus. The framework also enables the use of smaller LLMs for value extraction with estimable accuracy loss compared to large LLMs. We evaluate the framework on 5 health-related Reddit communities. Discovered attributes achieved 61% agreement with human-identified attributes, close to the 62% agreement between independent annotators. In most cases, the algorithm converges to a stable attribute set in fewer than 10 iterations. Structural type assignment achieves 82% accuracy, and value extraction reaches an F1 score of 0.8 compared to human annotations. Across four LLM families, smaller instruction-tuned models show statistically significant improvements in extraction performance with model scale when evaluated against a high-capacity reference LLM, supporting informed accuracy-cost trade-offs. These results show that attributes comparable to those identified by humans can be discovered automatically, enabling the creation of high-quality structured datasets economically and at scale. By removing the need for predefined ontologies, iterative model-driven schema induction offers a practical and scalable foundation for mining thematic corpora.
Create a lesson
Related papers
Optimizing Effective Training Time for Large-Scale Recommendation Systems
Mingming Ding, Ruilin Chen, Yuzhen Huang et al.
AgentWebRec: Compact Evidence Fusion over the Agent Web for Personalized Recommendation
Haoran Qiang, Guannan Liu, Liang Zhang et al.
From Rules to Neural Graphs: Scalable Structured Prediction for Patent Prior Art Search
Nikolai Zenovkin, Sebastian Björkqvist
Do Multilingual Encoders Produce Language-Consistent Semantic IDs?
Abhinav Bohra, Anuj Bohra
RPTune: Learned Context Curation for LLM Catalog Search
Chuxuan Hu, Hejie Cui, Norman Huang et al.
CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG
Hyojeong Yun, Jueun Kim, Wook-Shin Han