Variation and Synthetic SpeechWe describe the approach to linguistic variation taken by the Motorola speech synthesizer. A pan-dialectal pronunciation dictionary is described, which serves as the training data for a neural…Corey Miller, Orhan Karaali, Noel Massey·Nov 17, 1997SaveLearn
Probabilistic Parsing Using Left Corner Language ModelsWe introduce a novel parser based on a probabilistic version of a left-corner parser. The left-corner strategy is attractive because rule probabilities can be conditioned on both top-down goals and…Christopher D. Manning, Bob Carpenter·Nov 17, 1997SaveLearn
Approximating Context-Free Grammars with a Finite-State CalculusAlthough adequate models of human language for syntactic analysis and semantic interpretation are of at least context-free complexity, for applications such as speech processing in which speed is…Edmund Grimley-Evans·Nov 11, 1997SaveLearn
Probabilistic Constraint Logic ProgrammingThis paper addresses two central problems for probabilistic processing models: parameter estimation from incomplete data and efficient retrieval of most probable analyses. These questions have been…Stefan Riezler·Nov 11, 1997SaveLearn
Probabilistic Event CategorizationThis paper describes the automation of a new text categorization task. The categories assigned in this task are more syntactically, semantically, and contextually complex than those typically…Janyce Wiebe, Rebecca Bruce, Lei Duan·Oct 31, 1997SaveLearn
A Corpus-Based Investigation of Definite Description UseWe present the results of a study of definite descriptions use in written texts aimed at assessing the feasibility of annotating corpora with information about definite description interpretation. We…Massimo Poesio, Renata Vieira·Oct 24, 1997SaveLearn
Learning Features that Predict Cue UsageOur goal is to identify the features that predict the occurrence and placement of discourse cues in tutorial explanations in order to aid in the automatic generation of explanations. Previous…Barbara Di Eugenio, Johanna D. Moore, Massimo Paolucci·Oct 22, 1997SaveLearn
Attaching Multiple Prepositional Phrases: Generalized Backed-off EstimationThere has recently been considerable interest in the use of lexically-based statistical techniques to resolve prepositional phrase attachments. To our knowledge, however, these investigations have…Paola Merlo, Matthew Crocker, Cathy Berthouzoz·Oct 16, 1997SaveLearn
Parsing syllables: modeling OT computationallyIn this paper, I propose to implement syllabification in OT as a parser. I propose several innovations that result in a finite and small candidate set. The candidate set problem is handled with…Michael Hammond·Oct 14, 1997SaveLearn
Disambiguating with Controlled DisjunctionsIn this paper, we propose a disambiguating technique called controlled disjunctions. This extension of the so-called named disjunctions relies on the relations existing between feature values…Philippe Blache·Oct 14, 1997SaveLearn
Tagging French Without Lexical Probabilities -- Combining Linguistic Knowledge And Statistical LearningThis paper explores morpho-syntactic ambiguities for French to develop a strategy for part-of-speech disambiguation that a) reflects the complexity of French as an inflected language, b) optimizes…Evelyne Tzoukermann, Dragomir R. Radev, William A. Gale·Oct 10, 1997SaveLearn
Use of Weighted Finite State Transducers in Part of Speech TaggingThis paper addresses issues in part of speech disambiguation using finite-state transducers and presents two main contributions to the field. One of them is the use of finite-state machines for part…Evelyne Tzoukermann, Dragomir R. Radev·Oct 10, 1997SaveLearn
Segmentation of Expository Texts by Hierarchical Agglomerative ClusteringWe propose a method for segmentation of expository texts based on hierarchical agglomerative clustering. The method uses paragraphs as the basic segments for identifying hierarchical discourse…Yaakov Yaari·Sep 26, 1997SaveLearn
Amalia -- A Unified Platform for Parsing and GenerationContemporary linguistic theories (in particular, HPSG) are declarative in nature: they specify constraints on permissible structures, not how such structures are to be computed. Grammars designed…Shuly Wintner, Evgeniy Gabrilovich, Nissim Francez·Sep 24, 1997SaveLearn
An Abstract Machine for Unification GrammarsThis work describes the design and implementation of an abstract machine, Amalia, for the linguistic formalism ALE, which is based on typed feature structures. This formalism is one of the most…Shuly Wintner·Sep 23, 1997SaveLearn
Using Single Layer Networks for Discrete, Sequential Data: An Example from Natural Language ProcessingA natural language parser which has been successfully implemented is described. This is a hybrid system, in which neural networks operate within a rule based framework. It can be accessed via telnet…Caroline Lyon, Ray Frank·Sep 23, 1997SaveLearn
Off-line Parsability and the Well-foundedness of SubsumptionTyped feature structures are used extensively for the specification of linguistic information in many formalisms. The subsumption relation orders TFSs by their information content. We prove that…Shuly Wintner, Nissim Francez·Sep 23, 1997SaveLearn
Message-Passing Protocols for Real-World Parsing -- An Object-Oriented Model and its Preliminary EvaluationWe argue for a performance-based design of natural language grammars and their associated parsers in order to meet the constraints imposed by real-world NLP. Our approach incorporates declarative and…Udo Hahn, Peter Neuhaus, Norbert Broeker·Sep 23, 1997SaveLearn
Evaluating Parsing Schemes with Entropy IndicatorsThis paper introduces an objective metric for evaluating a parsing scheme. It is based on Shannon's original work with letter sequences, which can be extended to part-of-speech tag sequences. It…Caroline Lyon, Stephen Brown·Sep 22, 1997SaveLearn
Semantic Similarity Based on Corpus Statistics and Lexical TaxonomyThis paper presents a new approach for measuring semantic similarity/distance between words and concepts. It combines a lexical taxonomy structure with corpus statistical information so that the…Jay J. Jiang, David W. Conrath·Sep 20, 1997SaveLearn
Using WordNet to Complement Training Information in Text CategorizationAutomatic Text Categorization (TC) is a complex and useful task for many natural language applications, and is usually performed through the use of a set of manually classified documents, a training…Manuel de Buenaga Rodriguez, Jose Maria Gomez Hidalgo, Belen Diaz Agudo·Sep 17, 1997SaveLearn
Semantic Processing of Out-Of-Vocabulary Words in a Spoken Dialogue SystemOne of the most important causes of failure in spoken dialogue systems is usually neglected: the problem of words that are not covered by the system's vocabulary (out-of-vocabulary or OOV words).…Manuela Boros, Maria Aretoulaki, Florian Gallwitz et al.·Sep 17, 1997SaveLearn
A generation algorithm for f-structure representationsThis paper shows that previously reported generation algorithms run into problems when dealing with f-structure representations. A generation algorithm that is suitable for this type of…Toni Tuells·Sep 17, 1997SaveLearn
Combining Multiple Methods for the Automatic Construction of Multilingual WordNetsThis paper explores the automatic construction of a multilingual Lexical Knowledge Base from preexisting lexical resources. First, a set of automatic and complementary techniques for linking Spanish…Jordi Atserias, Salvador Climent, Xavier Farreres et al.·Sep 15, 1997SaveLearn
Integrating a Lexical Database and a Training Collection for Text CategorizationAutomatic text categorization is a complex and useful task for many natural language processing applications. Recent approaches to text categorization focus more on algorithms than on resources…Jose Maria Gomez Hidalgo, Manuel de Buenaga Rodriguez·Sep 15, 1997SaveLearn