Methodology · research-signals-v1
Research signals
ArcXiv assesses scientific promise from the sources available to it. Profiles help discovery; they are not peer review and do not establish that a paper’s claims are true.
What we assess
Every newly ingested paper receives a preliminary assessment from metadata and abstract. Papers selected by readers, notebooks, category leaders, comparisons, or digests may receive a full-paper assessment. Unsupported dimensions remain empty.
Research signals
- Impact
- The scale of change the work could create if its claims hold.
- Significance
- How important the addressed problem and contribution are to the field.
- Rigor
- The strength and care of the methods, analysis, and argument.
- Novelty
- How distinctly the work departs from existing ideas or methods.
- Clarity
- How precisely the paper communicates its question, method, and claims.
- Difficulty
- The background and effort a reader needs to understand the work.
- Surprisingness
- How far the findings or approach depart from reasonable expectations.
- Reproducibility
- How readily another group could repeat the work from the available detail.
- Translational potential
- How plausibly the work could influence practice or downstream research.
- Evidence strength
- How directly and convincingly the reported evidence supports the claims.
- Generalisability
- How well the claims may transfer beyond the studied setting.
- Interdisciplinarity
- How substantively the work connects questions or methods across fields.
- Refutation value
- How valuable the work would be if it overturned an accepted claim.
- Replication value
- How much independent repetition would strengthen the research record.
- Resource intensity
- The compute, data, equipment, time, or access needed to reproduce the work.
- Foundationality
- How useful the work may be as a basis for later research.
Public quality
The public quality value is deterministic and viewer-independent. Available weights are normalized: Impact 20%, Significance 15%, Rigor 20%, Novelty 15%, Evidence strength 10%, Reproducibility 5%, Generalisability 5%, Foundationality 5%, Clarity 5%. Impact, significance, rigor, novelty, and clarity must be present before a profile becomes public. Preliminary confidence cannot exceed 65%.
Category standing
Papers are compared within their primary arXiv category over 7 days, 30 days, 12 months, and all assessed papers. No percentile appears before 20 valid papers form a cohort. The feed uses the rolling 30-day standing. Popularity never changes public quality.
Close comparisons
Only the leading 10% of a cohort, plus papers within 0.35 quality points of its boundary, enter close comparison. Each selected paper meets 8–12 nearby peers. Bradley–Terry strength contributes 30% of ordering inside that set and deterministic quality 70%. Comparisons cannot move a paper across the set boundary.
Source handling
Full-paper jobs stream a PDF through the private arXiv gateway into encrypted temporary task storage. ArcXiv stores concise generated reasons and structural locators, then deletes the PDF and extracted body. PDF bytes, extracted text, and long excerpts are never retained.
Prompt templates
Preliminary: identify the paper type; assess each signal independently from title, categories, and abstract; use null when the source cannot support a judgment; separate proposed method from reported evidence; return strict structured data.
Full paper: use page-bounded, section-aware text; distinguish reported claims from verified facts; attach structural locators; never claim that an experiment was independently performed; return strict structured data.
Comparison: compare scientific promise inside a shared category without rank or popularity; allow a tie when evidence is insufficient; return a short reason and confidence. Reversed-order samples are used to monitor position bias.
Coverage
| Category | Preliminary | Full paper | Total |
|---|---|---|---|
| math-ph | 1411 | 0 | 1411 |
| hep-th | 1312 | 0 | 1312 |
| physics.gen-ph | 190 | 0 | 190 |
| cs.LG | 179 | 0 | 179 |
| cs.CV | 148 | 0 | 148 |
| cs.AI | 142 | 0 | 142 |
| quant-ph | 122 | 0 | 122 |
| physics.optics | 118 | 0 | 118 |
| cs.CL | 112 | 0 | 112 |
| physics.bio-ph | 101 | 0 | 101 |
| math.CO | 92 | 0 | 92 |
| math.AG | 90 | 0 | 90 |
| physics.flu-dyn | 88 | 0 | 88 |
| cs.RO | 79 | 0 | 79 |
| physics.chem-ph | 78 | 0 | 78 |
| math.AP | 76 | 0 | 76 |
| physics.atom-ph | 67 | 0 | 67 |
| physics.plasm-ph | 66 | 0 | 66 |
| math.DG | 65 | 0 | 65 |
| physics.comp-ph | 62 | 0 | 62 |
New-paper assessments run continuously after ingestion. Category standings update after valid profiles complete; close-comparison refinement runs on a scheduled cadence.
Versions and limitations
Provider, model, prompt, paper version, input basis, and assessment time are recorded. A new paper version creates a new assessment while history is retained. Results can reflect uncertainty, incomplete reporting, extraction limits, disciplinary differences, and model or prompt drift. Comparative validation claims are not published without reproducible data, sample sizes, confidence intervals, and commissioning disclosure.
Editorial accountability
Research signals are AI-assisted discovery aids, not peer review or author statements. The editorial and AI policy explains review status, content labels, and corrections. Report a paper-specific error through the published contact channel with the URL, disputed text, and a primary source.