Skip to content

Automatic Extraction of Tagset Mappings from Parallel-Annotated Corpora

John Hughes, Clive Souter, Eric Atwell

cmp-lgarXiv:cmp-lg/9506006

Abstract

This paper describes some of the recent work of project AMALGAM (automatic mapping among lexico-grammatical annotation models). We are investigating ways to map between the leading corpus annotation schemes in order to improve their resuability. Collation of all the included corpora into a single large annotated corpus will provide a more detailed language model to be developed for tasks such as speech and handwriting recognition. In particular, we focus here on a method of extracting mappings from corpora that have been annotated according to more than one annotation scheme.

Create a lesson