Initial Experiences Re-Exporting Duplicate and Similarity Computation with an OAI-PMH aggregator
Terry L. Harrison, Aravind Elango, Johan Bollen, Michael Nelson
Abstract
The proliferation of the Open Archive Initiative Protocol for Metadata Harvesting (OAI-PMH) has resulted in the creation of a large number of service providers, all harvesting from either data providers or aggregators. If data were available regarding the similarity of metadata records, service providers could track redundant records across harvests from multiple sources as well as provide additional end-user services. Due to the large number of metadata formats and the diverse mapping strategies employed by data providers, similarity calculation requirements necessitate the use of information retrieval strategies. We describe an OAI-PMH aggregator implementation that uses the optional ``<about>'' container to re-export the results of similarity calculations. Metadata records (3751) were harvested from a NASA data provider and similarities for the records were computed. The results were useful for detecting duplicates, similarities and metadata errors.
Create a lesson
Related papers
greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI
Justin Payan, Bálint Gyevnár, Atoosa Kasirzadeh et al.
Shifting Research Funding Priorities under Geopolitical Pressure: Evidence from Estonia
Yunfeng Gao, Yang Ding
Geospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines
Daniel Ebanks, Devika Jain
Quantifying the impact of clinical-academic collaborations
Mohamad Zeina, Nick McNally, Karl S. Peggs et al.
Testing Our Foundations: Citation Trends, Errors, and Emerging Hallucinations in the Computing Education Literature
Paul Denny, Gweneth Barbre, Musa Blake et al.
Toward non-textual representation of social anthropology: Modeling cultures as knowledge graphs
Manolis Peponakis, Sarantos Kapidakis, Martin Doerr et al.