A Linear Observed Time Statistical Parser Based on Maximum Entropy ModelsThis paper presents a statistical parser for natural language that obtains a parsing accuracy---roughly 87% precision and 86% recall---which surpasses the best previously published results on the…Adwait Ratnaparkhi·Jun 11, 1997SaveLearn
A Corpus-Based Approach for Building Semantic LexiconsSemantic knowledge can be a great asset to natural language processing systems, but it is usually hand-coded for each application. Although some semantic information is available in general-purpose…Ellen Riloff, Jessica Shepherd·Jun 10, 1997SaveLearn
Probabilistic Coreference in Information ExtractionCertain applications require that the output of an information extraction system be probabilistic, so that a downstream system can reliably fuse the output with possibly contradictory information…Andrew Kehler·Jun 10, 1997SaveLearn
Applying Reliability Metrics to Co-Reference AnnotationStudies of the contextual and linguistic factors that constrain discourse phenomena such as reference are coming to depend increasingly on annotated language corpora. In preparing the corpora, it is…Rebecca J. Passonneau·Jun 10, 1997SaveLearn
Exemplar-Based Word Sense Disambiguation: Some Recent ImprovementsIn this paper, we report recent improvements to the exemplar-based learning approach for word sense disambiguation that have achieved higher disambiguation accuracy. By using a larger value of k,…Hwee Tou Ng·Jun 10, 1997SaveLearn
Library of Practical Abstractions, Release 1.2The library of practical abstractions (LIBPA) provides efficient implementations of conceptually simple abstractions, in the C programming language. We believe that the best library code is…Eric Sven Ristad, Peter N. Yianilos·Jun 10, 1997SaveLearn
Distinguishing Word Senses in Untagged TextThis paper describes an experimental comparison of three unsupervised learning algorithms that distinguish the sense of an ambiguous word in untagged text. The methods described in this paper,…Ted Pedersen, Rebecca Bruce·Jun 9, 1997SaveLearn
Aggregate and mixed-order Markov models for statistical language processingWe consider the use of language models whose size and accuracy are intermediate between different order n-gram models. Two types of models are studied in particular. Aggregate Markov models are…Lawrence Saul, Fernando Pereira·Jun 9, 1997SaveLearn
Mistake-Driven Learning in Text CategorizationLearning problems in the text processing domain often map the text to a space whose dimensions are the measured features of the text, e.g., its words. Three characteristic properties of this domain…Ido Dagan, Yael Karov, Dan Roth·Jun 9, 1997SaveLearn
Comparing a Linguistic and a Stochastic TaggerConcerning different approaches to automatic PoS tagging: EngCG-2, a constraint-based morphological tagger, is compared in a double-blind test with a state-of-the-art statistical tagger on a common…Christer Samuelsson, Atro Voutilainen·Jun 7, 1997SaveLearn
Three New Probabilistic Models for Dependency Parsing: An ExplorationAfter presenting a novel O(n3) parsing algorithm for dependency grammar, we develop three contrasting ways to stochasticize it. We propose (a) a lexical affinity model where words struggle to modify…Jason Eisner·Jun 6, 1997SaveLearn
An Empirical Comparison of Probability Models for Dependency GrammarThis technical report is an appendix to Eisner (1996): it gives superior experimental results that were reported only in the talk version of that paper. Eisner (1996) trained three probability models…Jason Eisner·Jun 6, 1997SaveLearn
Learning Parse and Translation Decisions From Examples With Rich ContextWe present a knowledge and context-based system for parsing and translating natural language and evaluate it on sentences from the Wall Street Journal. Applying machine learning techniques, the…Ulf Hermjakob, Raymond J. Mooney·Jun 5, 1997SaveLearn
Assigning Grammatical Relations with a Back-off ModelThis paper presents a corpus-based method to assign grammatical subject/object relations to ambiguous German constructs. It makes use of an unsupervised learning procedure to collect training and…Erika F. de Lima·Jun 4, 1997SaveLearn
Sense Tagging: Semantic Tagging with a LexiconSense tagging, the automatic assignment of the appropriate sense from some lexicon to each of the words in a text, is a specialised instance of the general problem of semantic tagging by category or…Yorick Wilks, Mark Stevenson·May 29, 1997SaveLearn
Translation Methodology in the Spoken Language Translator: An EvaluationIn this paper we describe how the translation methodology adopted for the Spoken Language Translator (SLT) addresses the characteristics of the speech translation task in a context where it is…David Carter, Ralph Becket, Manny Rayner et al.·May 27, 1997SaveLearn
Incorporating POS Tagging into Language ModelingLanguage models for speech recognition tend to concentrate solely on recognizing the words that were spoken. In this paper, we redefine the speech recognition problem so that its goal is to find both…Peter A. Heeman, James F. Allen·May 22, 1997SaveLearn
FASTUS: A Cascaded Finite-State Transducer for Extracting Information from Natural-Language TextFASTUS is a system for extracting information from natural language text for entry into a database and for other applications. It works essentially as a cascaded, nondeterministic finite-state…Jerry R. Hobbs, Douglas Appelt, John Bear et al.·May 20, 1997SaveLearn
A Comparative Study of the Application of Different Learning Techniques to Natural Language InterfacesIn this paper we present first results from a comparative study. Its aim is to test the feasibility of different inductive learning techniques to perform the automatic acquisition of linguistic…Werner Winiwarter, Yahiko Kambayashi·May 16, 1997SaveLearn
A Lexicon for Underspecified Semantic TaggingThe paper defends the notion that semantic tagging should be viewed as more than disambiguation between senses. Instead, semantic tagging should be a first step in the interpretation process by…Paul Buitelaar·May 14, 1997SaveLearn
Memory-Based Learning: Using Similarity for SmoothingThis paper analyses the relation between the use of similarity in Memory-Based Learning and the notion of backed-off smoothing in statistical language modeling. We show that the two approaches are…Jakub Zavrel, Walter Daelemans·May 12, 1997SaveLearn
Charts, Interaction-Free Grammars, and the Compact Representation of AmbiguityRecently researchers working in the LFG framework have proposed algorithms for taking advantage of the implicit context-free components of a unification grammar [Maxwell 96]. This paper clarifies the…Marc Dymetman·May 12, 1997SaveLearn
The TreeBanker: a Tool for Supervised Training of Parsed CorporaI describe the TreeBanker, a graphical tool for the supervised training involved in domain customization of the disambiguation component of a speech- or language-understanding system. The TreeBanker…David Carter·May 7, 1997SaveLearn
Recycling Lingware in a Multilingual MT SystemWe describe two methods relevant to multi-lingual machine translation systems, which can be used to port linguistic data (grammars, lexicons and transfer rules) between systems used for processing…Manny Rayner, David Carter, Ivan Bretan et al.·May 7, 1997SaveLearn
Quantitative Constraint Logic Programming for Weighted Grammar ApplicationsConstraint logic grammars provide a powerful formalism for expressing complex logical descriptions of natural language phenomena in exact terms. Describing some of these phenomena may, however,…Stefan Riezler·May 6, 1997SaveLearn