Automatically Creating Bilingual Lexicons for Machine Translation from Bilingual TextA method is presented for automatically augmenting the bilingual lexicon of an existing Machine Translation system, by extracting bilingual entries from aligned bilingual text. The proposed method…Davide Turcato·Jul 20, 1998SaveLearn
A Linguistically Interpreted Corpus of German Newspaper TextIn this paper, we report on the development of an annotation scheme and annotation tools for unrestricted German text. Our representation format is based on argument structure, but also permits the…Wojciech Skut, Thorsten Brants, Brigitte Krenn et al.·Jul 17, 1998SaveLearn
Chunk Tagger - Statistical Recognition of Noun PhrasesWe describe a stochastic approach to partial parsing, i.e., the recognition of syntactic structures of limited depth. The technique utilises Markov Models, but goes beyond usual bracketing…Wojciech Skut, Thorsten Brants·Jul 17, 1998SaveLearn
A Maximum-Entropy Partial Parser for Unrestricted TextThis paper describes a partial parser that assigns syntactic structures to sequences of part-of-speech tags. The program uses the maximum entropy parameter estimation method, which allows a flexible…Wojciech Skut, Thorsten Brants·Jul 17, 1998SaveLearn
Graph Interpolation Grammars as Context-Free AutomataA derivation step in a Graph Interpolation Grammar has the effect of scanning an input token. This feature, which aims at emulating the incrementality of the natural parser, restricts the formal…John Larcheveque·Jul 17, 1998SaveLearn
Word Clustering and Disambiguation Based on Co-occurrence DataWe address the problem of clustering words (or constructing a thesaurus) based on co-occurrence data, and using the acquired word classes to improve the accuracy of syntactic disambiguation. We view…Hang Li, Naoki Abe·Jul 17, 1998SaveLearn
Centering in Dynamic SemanticsCentering theory posits a discourse center, a distinguished discourse entity that is the topic of a discourse. A simplified version of this theory is developed in a Dynamic Semantics framework. In…Daniel Hardt·Jul 14, 1998SaveLearn
The Role of Verbs in Document AnalysisWe present results of two methods for assessing the event profile of news articles as a function of verb type. The unique contribution of this research is the focus on the role of verbs, rather than…Judith L. Klavans, Min-Yen Kan·Jul 13, 1998SaveLearn
Evaluating a Focus-Based Approach to Anaphora ResolutionWe present an approach to anaphora resolution based on a focusing algorithm, and implemented within an existing MUC (Message Understanding Conference) Information Extraction system, allowing…Saliha Azzam, Kevin Humphreys, Robert Gaizauskas·Jul 6, 1998SaveLearn
Textual Economy through Close Coupling of Syntax and SemanticsWe focus on the production of efficient descriptions of objects, actions and events. We define a type of efficiency, textual economy, that exploits the hearer's recognition of inferential links…Matthew Stone, Bonnie Webber·Jun 29, 1998SaveLearn
An Empirical Investigation of Proposals in Collaborative DialoguesWe describe a corpus-based investigation of proposals in dialogue. First, we describe our DRI compliant coding scheme and report our inter-coder reliability results. Next, we test several hypotheses…Barbara Di Eugenio, Pamela W. Jordan, Johanna D. Moore et al.·Jun 25, 1998SaveLearn
Never Look Back: An Alternative to CenteringI propose a model for determining the hearer's attentional state which depends solely on a list of salient discourse entities (S-list). The ordering among the elements of the S-list covers also…Michael Strube·Jun 25, 1998SaveLearn
Anchoring a Lexicalized Tree-Adjoining Grammar for DiscourseWe here explore a ``fully'' lexicalized Tree-Adjoining Grammar for discourse that takes the basic elements of a (monologic) discourse to be not simply clauses, but larger structures that are…Bonnie Lynn Webber, Aravind K. Joshi·Jun 24, 1998SaveLearn
Using WordNet for Building WordNetsThis paper summarises a set of methodologies and techniques for the fast construction of multilingual WordNets. The English WordNet is used in this approach as a backbone for Catalan and Spanish…Xavier Farreres, German Rigau, Horacio Rodriguez·Jun 23, 1998SaveLearn
Building Accurate Semantic Taxonomies from Monolingual MRDsThis paper presents a method that combines a set of unsupervised algorithms in order to accurately build large taxonomies from any machine-readable dictionary (MRD). Our aim is to profit from…German Rigau, Horacio Rodriguez, Eneko Agirre·Jun 23, 1998SaveLearn
Word Sense Disambiguation using Optimised Combinations of Knowledge SourcesWord sense disambiguation algorithms, with few exceptions, have made use of only one lexical knowledge source. We describe a system which performs unrestricted word sense disambiguation (on all…Yorick Wilks, Mark Stevenson·Jun 22, 1998SaveLearn
Can Subcategorisation Probabilities Help a Statistical Parser?Research into the automatic acquisition of lexical information from corpora is starting to produce large-scale computational lexicons containing data on the relative frequencies of subcategorisation…John Carroll, Guido Minnen, Ted Briscoe·Jun 21, 1998SaveLearn
Bayesian Stratified Sampling to Assess Corpus UtilityThis paper describes a method for asking statistical questions about a large text corpus. We exemplify the method by addressing the question, "What percentage of Federal Register documents are…Judith Hochberg, Clint Scovel, Timothy Thomas et al.·Jun 19, 1998SaveLearn
Towards a single proposal is spelling correctionThe study presented here relies on the integrated use of different kinds of knowledge in order to improve first-guess accuracy in non-word context-sensitive correction for general unrestricted texts.…E. Agirre, K. Gojenola, K. Sarasola·Jun 15, 1998SaveLearn
Methods and Tools for Building the Catalan WordNetIn this paper we introduce the methodology used and the basic phases we followed to develop the Catalan WordNet, and shich lexical resources have been employed in its building. This methodology, as…Laura Benitez, Sergi Cervell, Gerard Escudero et al.·Jun 11, 1998SaveLearn
Unlimited Vocabulary Grapheme to Phoneme Conversion for Korean TTSThis paper describes a grapheme-to-phoneme conversion method using phoneme connectivity and CCV conversion rules. The method consists of mainly four modules including morpheme normalization,…Byeongchang Kim, WonIl Lee, Geunbae Lee et al.·Jun 10, 1998SaveLearn
An Investigation of Transformation-Based Learning in DiscourseThis paper presents results from the first attempt to apply Transformation-Based Learning to a discourse-level Natural Language Processing task. To address two limitations of the standard algorithm,…Ken Samuel, Sandra Carberry, K. Vijay-Shanker·Jun 9, 1998SaveLearn
Dialogue Act Tagging with Transformation-Based LearningFor the task of recognizing dialogue acts, we are applying the Transformation-Based Learning (TBL) machine learning algorithm. To circumvent a sparse data problem, we extract values of well-motivated…Ken Samuel, Sandra Carberry, K. Vijay-Shanker·Jun 8, 1998SaveLearn
Eliminating deceptions and mistaken belief to infer conversational implicatureConversational implicatures are usually described as being licensed by the disobeying or flouting of some principle by the speaker in cooperative dialogue. However, such work has failed to…Mark Lee, Yorick Wilks·Jun 5, 1998SaveLearn
Rationality, Cooperation and Conversational ImplicatureConversational implicatures are usually described as being licensed by the disobeying or flouting of a Principle of Cooperation. However, the specification of this principle has proved…Mark Lee·Jun 5, 1998SaveLearn