Automatic Identification of Subjects for Textual Documents in Digital Libraries
Kuang-hua Chen
Abstract
The amount of electronic documents in the Internet grows very quickly. How to effectively identify subjects for documents becomes an important issue. In past, the researches focus on the behavior of nouns in documents. Although subjects are composed of nouns, the constituents that determine which nouns are subjects are not only nouns. Based on the assumption that texts are well-organized and event-driven, nouns and verbs together contribute the process of subject identification. This paper considers four factors: 1) word importance, 2) word frequency, 3) word co-occurrence, and 4) word distance and proposes a model to identify subjects for textual documents. The preliminary experiments show that the performance of the proposed model is close to that of human beings.
Create a lesson
Related papers
Gender and the Production of Research Impact
Sanger Wagner, Charles Rahal, Melinda C. Mills
A Comparative Evaluation of Digitization Pipelines for Historiographical Sources
Marina Gómez Rey, Patricia Callejo, Mario Muñoz-Organero et al.
COCI: Conference Organisers and Content Identifier
Angelo Salatino, Francesco Osborne, Alexis Vizcaino et al.
Towards a Definition of the Computational Architecture of Open Scholarly Infrastructures
Ivan Heibi, Mario Petrella, Angelo Di Iorio et al.
Beyond FAIR Data: Instrument Traces for Active and Autonomous Scientific Experimentation
Sergei V. Kalinin, Boris N. Slautin, Yu Liu et al.
Taxonomy-aware distances between scholarly topic profiles via an exact simplex embedding
Dmitry Gubanov, Alexander Chkhartishvili