Lazy Transformation-Based LearningWe introduce a significant improvement for a relatively new machine learning method called Transformation-Based Learning. By applying a Monte Carlo strategy to randomly sample from the space of…Ken Samuel·Jun 3, 1998SaveLearn
Computing Dialogue Acts from Features with Transformation-Based LearningTo interpret natural language at the discourse level, it is very useful to accurately recognize dialogue acts, such as SUGGEST, in identifying speaker intentions. Our research explores the utility of…Ken Samuel, Sandra Carberry, K. Vijay-Shanker·Jun 2, 1998SaveLearn
Learning Correlations between Linguistic Indicators and Semantic Constraints: Reuse of Context-Dependent Descriptions of EntitiesThis paper presents the results of a study on the semantic constraints imposed on lexical choice by certain contextual indicators. We show how such indicators are computed and how correlations…Dragomir R. Radev·May 31, 1998SaveLearn
Recognizing Syntactic Errors in the Writing of Second Language LearnersThis paper reports on the recognition component of an intelligent tutoring system that is designed to help foreign language speakers learn standard English. The system models the grammar of the…David A. Schneider, Kathleen F. McCoy·May 29, 1998SaveLearn
Automatic summarising: factors and directionsThis position paper suggests that progress with automatic summarising demands a better research methodology and a carefully focussed research strategy. In order to develop effective procedures it is…Karen Sparck Jones·May 29, 1998SaveLearn
Integrating Text Plans for Conciseness and CoherenceOur experience with a critiquing system shows that when the system detects problems with the user's performance, multiple critiques are often produced. Analysis of a corpus of actual critiques…Terrence Harvey, Sandra Carberry·May 28, 1998SaveLearn
Discovery of Linguistic Relations Using Lexical AttractionThis work has been motivated by two long term goals: to understand how humans learn language and to build programs that can understand language. Using a representation that makes the relevant…Deniz Yuret·May 27, 1998SaveLearn
A Descriptive Characterization of Tree-Adjoining Languages (Full Version)Since the early Sixties and Seventies it has been known that the regular and context-free languages are characterized by definability in the monadic second-order theory of certain structures. More…James Rogers·May 21, 1998SaveLearn
Parsing Inside-OutThe inside-outside probabilities are typically used for reestimating Probabilistic Context Free Grammars (PCFGs), just as the forward-backward probabilities are typically used for reestimating HMMs.…Joshua Goodman·May 19, 1998SaveLearn
The Proper Treatment of Optimality in Computational PhonologyThis paper presents a novel formalization of optimality theory. Unlike previous treatments of optimality in computational linguistics, starting with Ellison (1994), the new approach does not require…Lauri Karttunen·May 12, 1998SaveLearn
Word-to-Word Models of Translational EquivalenceParallel texts (bitexts) have properties that distinguish them from other kinds of parallel data. First, most words translate to only one other word. Second, bitext correspondence is noisy. This…I. Dan Melamed·May 11, 1998SaveLearn
Manual Annotation of Translational Equivalence: The Blinker ProjectBilingual annotators were paid to link roughly sixteen thousand corresponding words between on-line versions of the Bible in modern French and modern English. These annotations are freely available…I. Dan Melamed·May 11, 1998SaveLearn
Annotation Style Guide for the Blinker ProjectThis annotation style guide was created by and for the Blinker project at the University of Pennsylvania. The Blinker project was so named after the ``bilingual linker'' GUI, which was…I. Dan Melamed·May 8, 1998SaveLearn
Models of Co-occurrenceA model of co-occurrence in bitext is a boolean predicate that indicates whether a given pair of word tokens co-occur in corresponding regions of the bitext space. Co-occurrence is a precondition for…I. Dan Melamed·May 8, 1998SaveLearn
Group Theory and Grammatical DescriptionThis paper presents a model for linguistic description based on group theory. A grammar in this model, or "G-grammar", is a collection of lexical expressions which are products of logical…Marc Dymetman·May 7, 1998SaveLearn
Valence Induction with a Head-Lexicalized PCFGThis paper presents an experiment in learning valences (subcategorization frames) from a 50 million word text corpus, based on a lexicalized probabilistic context free grammar. Distributions are…Glenn Carroll, Mats Rooth·May 5, 1998SaveLearn
On the existence of certain total recursive functions in nontrivial axiom systems, IWe investigate the existence of a class of ZFC-provably total recursive unary functions, given certain constraints, and apply some of those results to show that, for Σ1-sound set theory,…N. C. A. da Costa, F. A. Doria·Apr 30, 1998SaveLearn
Corpus-Based Word Sense DisambiguationResolution of lexical ambiguity, commonly termed ``word sense disambiguation'', is expected to improve the analytical accuracy for tasks which are sensitive to lexical semantics. Such tasks…Atsushi Fujii·Apr 29, 1998SaveLearn
Treatment of Epsilon-Moves in Subset ConstructionThe paper discusses the problem of determinising finite-state automata containing large numbers of epsilon-moves. Experiments with finite-state approximations of natural language grammars often give…Gertjan van Noord·Apr 28, 1998SaveLearn
Graph Interpolation Grammars: a Rule-based Approach to the Incremental Parsing of Natural LanguagesGraph Interpolation Grammars are a declarative formalism with an operational semantics. Their goal is to emulate salient features of the human parser, and notably incrementality. The parsing process…John Larcheveque·Apr 2, 1998SaveLearn
Nymble: a High-Performance Learning Name-finderThis paper presents a statistical, learned approach to finding names and other non-recursive entities in text (as per the MUC-6 definition of the NE task), using a variant of the standard hidden…Daniel M. Bikel, Scott Miller, Richard Schwartz et al.·Mar 27, 1998SaveLearn
Time, Tense and Aspect in Natural Language Database InterfacesMost existing natural language database interfaces (NLDBs) were designed to be used with database systems that provide very limited facilities for manipulating time-dependent data, and they do not…I. Androutsopoulos, G. D. Ritchie, P. Thanisch·Mar 22, 1998SaveLearn
Automating Coreference: The Role of Annotated Training DataWe report here on a study of interannotator agreement in the coreference task as defined by the Message Understanding Conference (MUC-6 and MUC-7). Based on feedback from annotators, we clarified and…Lynette Hirschman, Patricia Robinson, John Burger et al.·Mar 2, 1998SaveLearn
A Hybrid Environment for Syntax-Semantic TaggingThe thesis describes the application of the relaxation labelling algorithm to NLP disambiguation. Language is modelled through context constraint inspired on Constraint Grammars. The constraints…Lluis Padro·Feb 11, 1998SaveLearn
Look-Back and Look-Ahead in the Conversion of Hidden Markov Models into Finite State TransducersThis paper describes the conversion of a Hidden Markov Model into a finite state transducer that closely approximates the behavior of the stochastic model. In some cases the transducer is equivalent…Andre Kempe·Feb 2, 1998SaveLearn