Semi-Automatic Indexing of Multilingual Documents
Ulrich Schiel, Ianna M. Sodre Ferreira de Souza, Edberto Ferneda
Abstract
With the growing significance of digital libraries and the Internet, more and more electronic texts become accessible to a wide and geographically disperse public. This requires adequate tools to facilitate indexing, storage, and retrieval of documents written in different languages. We present a method for semi-automatic indexing of electronic documents and construction of a multilingual thesaurus, which can be used for query formulation and information retrieval. We use special dictionaries and user interaction in order to solve ambiguities and find adequate canonical terms in the language and adequate abstract language-independent terms. The abstract thesaurus is updated incrementally by new indexed documents and is used to search document concerning terms in a query to the document base.
Create a lesson
Related papers
Gender and the Production of Research Impact
Sanger Wagner, Charles Rahal, Melinda C. Mills
A Comparative Evaluation of Digitization Pipelines for Historiographical Sources
Marina Gómez Rey, Patricia Callejo, Mario Muñoz-Organero et al.
COCI: Conference Organisers and Content Identifier
Angelo Salatino, Francesco Osborne, Alexis Vizcaino et al.
Towards a Definition of the Computational Architecture of Open Scholarly Infrastructures
Ivan Heibi, Mario Petrella, Angelo Di Iorio et al.
Beyond FAIR Data: Instrument Traces for Active and Autonomous Scientific Experimentation
Sergei V. Kalinin, Boris N. Slautin, Yu Liu et al.
Taxonomy-aware distances between scholarly topic profiles via an exact simplex embedding
Dmitry Gubanov, Alexander Chkhartishvili