Integrating Syntactic and Prosodic Information for the Efficient Detection of Empty CategoriesWe describe a number of experiments that demonstrate the usefulness of prosodic information for a processing module which parses spoken utterances with a feature-based grammar employing empty…Anton Batliner, Anke Feldhaus, Stefan Geissler et al.·Jul 2, 1996SaveLearn
Domain and Language Independent Feature Extraction for Statistical Text CategorizationA generic system for text categorization is presented which uses a representative text corpus to adapt the processing steps: feature extraction, dimension reduction, and classification. Feature…Thomas Bayer, Ingrid Renz, Michael Stein et al.·Jul 2, 1996SaveLearn
Inducing Constraint GrammarsConstraint Grammar rules are induced from corpora. A simple scheme based on local information, i.e., on lexical biases and next-neighbour contexts, extended through the use of barriers, reached 87.3…Christer Samuelsson, Pasi Tapanainen, Atro Voutilainen·Jul 1, 1996SaveLearn
GramCheck: A Grammar and Style CheckerThis paper presents a grammar and style checker demonstrator for Spanish and Greek native writers developed within the project GramCheck. Besides a brief grammar error typology for Spanish, a…Flora Ramírez Bustamante, Fernando Sánchez León·Jul 1, 1996SaveLearn
Integrating Multiple Knowledge Sources to Disambiguate Word Sense: An Exemplar-Based ApproachIn this paper, we present a new approach for word sense disambiguation (WSD) using an exemplar-based learning algorithm. This approach integrates a diverse set of knowledge sources to disambiguate…Hwee Tou Ng, Hian Beng Lee·Jun 29, 1996SaveLearn
Research on Architectures for Integrated Speech/Language Systems in VerbmobilThe German joint research project Verbmobil (VM) aims at the development of a speech to speech translation system. This paper reports on research done in our group which belongs to Verbmobil's…Günther Görz, Marcus Kesseler, Jörg Spilker et al.·Jun 25, 1996SaveLearn
Minimizing Manual Annotation Cost In Supervised Training From CorporaCorpus-based methods for natural language processing often use supervised training, requiring expensive manual annotation of training corpora. This paper investigates methods for reducing annotation…Sean P. Engelson, Ido Dagan·Jun 24, 1996SaveLearn
Directed ReplacementThis paper introduces to the finite-state calculus a family of directed replace operators. In contrast to the simple replace expression, UPPER -> LOWER, defined in Karttunen (ACL-95), the new…Lauri Karttunen·Jun 23, 1996SaveLearn
Linguistic Structure as Composition and PerturbationThis paper discusses the problem of learning language from unprocessed text and speech signals, concentrating on the problem of learning a lexicon. In particular, it argues for a representation of…Carl de Marcken·Jun 21, 1996SaveLearn
Maximizing Top-down Constraints for Unification-based SystemsA left-corner parsing algorithm with top-down filtering has been reported to show very efficient performance for unification-based systems. However, due to the nontermination of parsing with…Noriko Tomuro·Jun 20, 1996SaveLearn
An Efficient Compiler for Weighted Rewrite RulesContext-dependent rewrite rules are used in many areas of natural language and speech processing. Work in computational phonology has demonstrated that, given certain conditions, such rewrite rules…Mehryar Mohri, Richard Sproat·Jun 20, 1996SaveLearn
Two Sources of Control over the Generation of Software InstructionsThis paper presents an analysis conducted on a corpus of software instructions in French in order to establish whether task structure elements (the procedural representation of the users' tasks)…Anthony Hartley, Cecile Paris·Jun 19, 1996SaveLearn
A Data-Oriented Approach to Semantic InterpretationIn Data-Oriented Parsing (DOP), an annotated language corpus is used as a stochastic grammar. The most probable analysis of a new input sentence is constructed by combining sub-analyses from the…Rens Bod, Remko Bonnema, Remko Scha·Jun 18, 1996SaveLearn
A Robust System for Natural Spoken DialogueThis paper describes a system that leads us to believe in the feasibility of constructing natural spoken dialogue systems in task-oriented domains. It specifically addresses the issue of robust…James F. Allen, Bradford W. Miller, Eric K. Ringger et al.·Jun 18, 1996SaveLearn
Two Questions about Data-Oriented ParsingIn this paper I present ongoing work on the data-oriented parsing (DOP) model. In previous work, DOP was tested on a cleaned-up set of analyzed part-of-speech strings from the Penn Treebank,…Rens Bod·Jun 17, 1996SaveLearn
An Iterative Algorithm to Build Chinese Language ModelsWe present an iterative procedure to build a Chinese language model (LM). We segment Chinese text into words based on a word-based Chinese language model. However, the construction of a Chinese LM…Xiaoqiang Luo, Salim Roukos·Jun 17, 1996SaveLearn
Computing Optimal Descriptions for Optimality Theory Grammars with Context-Free Position StructuresThis paper describes an algorithm for computing optimal structural descriptions for Optimality Theory grammars with context-free position structures. This algorithm extends Tesar's dynamic…Bruce Tesar·Jun 17, 1996SaveLearn
Computational Complexity of Probabilistic Disambiguation by means of Tree-GrammarsThis paper studies the computational complexity of disambiguation under probabilistic tree-grammars and context-free grammars. It presents a proof that the following problems are NP-hard: computing…Khalil Sima'an·Jun 17, 1996SaveLearn
Compilation of Weighted Finite-State Transducers from Decision TreesWe report on a method for compiling decision trees into weighted finite-state transducers. The key assumptions are that the tree predictions specify how to rewrite symbols from an input string, and…Richard Sproat, Michael Riley·Jun 14, 1996SaveLearn
With raised eyebrows or the eyebrows raised ? A Neural Network Approach to Grammar Checking for DefinitenessIn this paper, we use a feature model of the semantics of plural determiners to present an approach to grammar checking for definiteness. Using neural network techniques, a semantics -- morphological…Gabriele Scheler·Jun 14, 1996SaveLearn
A Probabilistic Disambiguation Method Based on Psycholinguistic PrinciplesWe address the problem of structural disambiguation in syntactic parsing. In psycholinguistics, a number of principles of disambiguation have been proposed, notably the Lexical Preference Rule (LPR),…Hang Li·Jun 13, 1996SaveLearn
Relating Turing's Formula and Zipf's LawAn asymptote is derived from Turing's local reestimation formula for population frequencies, and a local reestimation formula is derived from Zipf's law for the asymptotic behavior of…Christer Samuelsson·Jun 11, 1996SaveLearn
Stabilizing the Richardson Algorithm by Controlling ChaosBy viewing the operations of the Richardson purification algorithm as a discrete time dynamical process, we propose a method to overcome the instability of the algorithm by controlling chaos. We…Song He·Jun 11, 1996SaveLearn
Building Probabilistic Models for Natural LanguageIn this thesis, we investigate three problems involving the probabilistic modeling of language: smoothing n-gram models, statistical grammar induction, and bilingual sentence alignment. These three…Stanley F. Chen·Jun 11, 1996SaveLearn
An Efficient Inductive Unsupervised Semantic TaggerWe report our development of a simple but fast and efficient inductive unsupervised semantic tagger for Chinese words. A POS hand-tagged corpus of 348,000 words is used. The corpus is being tagged in…K T Lua·Jun 11, 1996SaveLearn