Learning to Paraphrase: An Unsupervised Approach Using Multiple-Sequence Alignment
Regina Barzilay, Lillian Lee
Abstract
We address the text-to-text generation problem of sentence-level paraphrasing -- a phenomenon distinct from and more difficult than word- or phrase-level paraphrasing. Our approach applies multiple-sequence alignment to sentences gathered from unannotated comparable corpora: it learns a set of paraphrasing patterns represented by word lattice pairs and automatically determines how to apply these patterns to rewrite new sentences. The results of our evaluation experiments show that the system derives accurate paraphrases, outperforming baseline systems.
Create a lesson
Related papers
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Yan Yu, Zhengxi Lu, Yizhou Liu et al.
Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Sarah Wyer, Sue Black, Noura Al Moubayed
dQwen3.5: Hybrid-Attention Diffusion Language Models
Anton Xue, Litu Rout, Aditya Akella et al.
On-Demand Attention: Language Models Know When to Recall
Haibo Feng, Ruiqi Liang, Hanyang Peng et al.
Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol
Levent Bulut
HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication
Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi et al.