Rerendering Semantic Ontologies: Automatic Extensions to UMLS through Corpus Analytics
J. Pustejovsky, A. Rumshisky, J. Castano
Abstract
In this paper, we discuss the utility and deficiencies of existing ontology resources for a number of language processing applications. We describe a technique for increasing the semantic type coverage of a specific ontology, the National Library of Medicine's UMLS, with the use of robust finite state methods used in conjunction with large-scale corpus analytics of the domain corpus. We call this technique "semantic rerendering" of the ontology. This research has been done in the context of Medstract, a joint Brandeis-Tufts effort aimed at developing tools for analyzing biomedical language (i.e., Medline), as well as creating targeted databases of bio-entities, biological relations, and pathway data for biological researchers. Motivating the current research is the need to have robust and reliable semantic typing of syntactic elements in the Medline corpus, in order to improve the overall performance of the information extraction applications mentioned above.
Create a lesson
Related papers
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
Yan Yu, Zhengxi Lu, Yizhou Liu et al.
Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Sarah Wyer, Sue Black, Noura Al Moubayed
dQwen3.5: Hybrid-Attention Diffusion Language Models
Anton Xue, Litu Rout, Aditya Akella et al.
On-Demand Attention: Language Models Know When to Recall
Haibo Feng, Ruiqi Liang, Hanyang Peng et al.
Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol
Levent Bulut
HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication
Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi et al.