Plagiarism Detection in arXiv
Daria Sorokina, Johannes Gehrke, Simeon Warner, Paul Ginsparg
Abstract
We describe a large-scale application of methods for finding plagiarism in research document collections. The methods are applied to a collection of 284,834 documents collected by arXiv.org over a 14 year period, covering a few different research disciplines. The methodology efficiently detects a variety of problematic author behaviors, and heuristics are developed to reduce the number of false positives. The methods are also efficient enough to implement as a real-time submission screen for a collection many times larger.
Create a lesson
Related papers
Resolution limits for process comparison from event data
Antony R. Lee, Peter Tiňo, Iain B. Styles
A Functional Pilot for Certified Freshness-Aware Semantic--Spatial Range Retrieval
Taimoor Ahmad
Efficiently Linking Unstructured Data for Multi-step Reasoning
Jiaming Liang, Haydn Jones, Jacob R. Gardner et al.
Distribution-Aware Distributed Database Testing (Extended Version)
Zhou Zhou, Si Liu, Hengfeng Wei et al.
Linking Speakers of the German Parliament to Wikidata: Scope and Coverage of Metadata
Thomas Haider, Arne Cypionka, Maximilian Teich
How Can We Shrink the Family of Test Databases? Query Containment with Nulls and Comparisons
Helen Sternbach, Sara Cohen