Development of a Spanish Version of the Xerox Tagger
Fernando Sánchez León, Amalio F. Nieto Serrano
Abstract
This paper describes work performed withing the CRATER ( Corpus Resources And Terminology Ext Raction, MLAP-93/20) project, funded by the Commission of the European Communities. In particular, it addresses the issue of adapting the Xerox Tagger to Spanish in order to tag the Spanish version of the ITU (International Telecommunications Union) corpus. The model implemented by this tagger is briefly presented along with some modifications performed on it in order to use some parameters not probabilistically estimated. Initial decisions, like the tagset, the lexicon and the training corpus are also discussed. Finally, results are presented and the benefits of the mixed model justified.
Create a lesson
Related papers
A Memory-Based Approach to Learning Shallow Natural Language Patterns
Shlomo Argamon, Ido Dagan, Yuval Krymolowski
A Comparison of WordNet and Roget's Taxonomy for Measuring Semantic Similarity
Michael Mc Hale
Some Ontological Principles for Designing Upper Level Lexical Resources
Nicola Guarino
Towards an implementable dependency grammar
Timo Jarvinen, Pasi Tapanainen
A Variant of Earley Parsing
Mark-Jan Nederhof, Giorgio Satta
Segregatory Coordination and Ellipsis in Text Generation
James Shaw