Combining semantic and syntactic structure for language modeling
Rens Bod
Abstract
Structured language models for speech recognition have been shown to remedy the weaknesses of n-gram models. All current structured language models are, however, limited in that they do not take into account dependencies between non-headwords. We show that non-headword dependencies contribute to significantly improved word error rate, and that a data-oriented parsing model trained on semantically and syntactically annotated data can exploit these dependencies. This paper also contains the first DOP model trained by means of a maximum likelihood reestimation procedure, which solves some of the theoretical shortcomings of previous DOP models.
Create a lesson
Related papers
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Yan Yu, Zhengxi Lu, Yizhou Liu et al.
Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Sarah Wyer, Sue Black, Noura Al Moubayed
dQwen3.5: Hybrid-Attention Diffusion Language Models
Anton Xue, Litu Rout, Aditya Akella et al.
On-Demand Attention: Language Models Know When to Recall
Haibo Feng, Ruiqi Liang, Hanyang Peng et al.
Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol
Levent Bulut
HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication
Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi et al.