Pith. sign in

REVIEW 15 cited by

SPECTER: Document-level Representation Learning using Citation-informed Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.07180 v4 pith:WDYBXQAR submitted 2020-04-15 cs.CL

classification cs.CL
keywords document-levellanguagemodelsspecterrepresentationapplicationsbenchmarkcitation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Representation learning is a critical ingredient for natural language processing systems. Recent Transformer language models like BERT learn powerful textual representations, but these models are targeted towards token- and sentence-level training objectives and do not leverage information on inter-document relatedness, which limits their document-level representation power. For applications on scientific documents, such as classification and recommendation, the embeddings power strong performance on end tasks. We propose SPECTER, a new method to generate document-level embedding of scientific documents based on pretraining a Transformer language model on a powerful signal of document-level relatedness: the citation graph. Unlike existing pretrained language models, SPECTER can be easily applied to downstream applications without task-specific fine-tuning. Additionally, to encourage further research on document-level models, we introduce SciDocs, a new evaluation benchmark consisting of seven document-level tasks ranging from citation prediction, to document classification and recommendation. We show that SPECTER outperforms a variety of competitive baselines on the benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 274 citations worldwide. Full citation record

  1. ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Conference accept/reject outcomes yield 15 operational ideation patterns that, as an LLM skill suite, improve automated-judged research-proposal quality over no-skill and generic-skill baselines.

  2. Mutual Linearity in and out of Stationarity for Markov Jump Processes: A Trajectory-Based Approach

    cond-mat.stat-mech 2026-04 unverdicted novelty 7.0 of 10

    Trajectory-level linear response yields mutual linearity of observables under single-edge rate perturbation for Markov jump processes, including non-stationary state and counting observables.

  3. BitNet Text Embeddings

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    BITEMBED trains 1.58-bit ternary-weight LLM embedders with contrastive pre-training, supervised distillation, and multi-precision output training, matching FP16 teachers within ~0.6 MMTEB points at ~2x CPU speed.

  4. Col-Bandit: Query-Time Top-$K$ Estimation for Late-Interaction Retrieval

    cs.IR 2026-02 conditional novelty 6.0 of 10

    An adaptive confidence-bound cell-pruning method recovers the exhaustive MaxSim top-K with roughly one-quarter to one-third of the compute on BEIR and REAL-MM-RAG, at the price of a calibrated rather than certified de...

  5. From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A weak-supervision pipeline using LLM-extracted concepts and rank fusion creates silver-standard patent-to-SDG labels that recover known citation-derived associations and show high network modularity.

  6. SPAR: Scholar Paper Retrieval with LLM-based Agents for Enhanced Academic Search

    cs.IR 2025-07 conditional novelty 6.0 of 10

    An agentic academic search system using LLM agents for query understanding, citation-chain exploration, and reranking outperforms prior retrieval baselines on two benchmarks.

  7. Four Shades of Life Sciences: A Dataset for Disinformation Detection in the Life Sciences

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Introduces FSoLS, a four-class labeled corpus of 2,603 full-text life-science articles, and benchmarks language models that classify disinformative texts with up to 98% F1.

  8. MIR: Methodology Inspiration Retrieval for Scientific Research Problems

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A new dataset and MAG-guided triplet-loss fine-tuning, plus LLM reranking, improves retrieval of methodologically inspirational papers for research proposals by several points over strong baselines.

  9. Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs

    cs.CL 2025-10 conditional novelty 5.0 of 10

    An LLM-based pipeline extracts technology triples from arXiv and patent full text, tracks rising topic co-occurrence, and labels retrieval-augmented generation and conversational agents as emerging transformative tech...

  10. A Comparative Study of Specialized LLMs as Dense Retrievers

    cs.IR 2025-07 conditional novelty 5.0 of 10

    Specialized Qwen2.5 7B models differ in dense retrieval quality: math and long-reasoning variants degrade performance, while coder and vision-language variants improve zero-shot text and code retrieval.

  11. Extracting Information About Publication Venues Using Citation-Informed Transformers

    cs.DL 2025-06 conditional novelty 5.0 of 10

    SPECTER paper embeddings show that some CS venues are nearly indistinguishable and that several venue pairs converged between 2015 and 2023.

  12. exHarmony: Authorship and Citations for Benchmarking the Reviewer Assignment Problem

    cs.IR 2025-02 conditional novelty 5.0 of 10

    exHarmony is a public benchmark that turns reviewer assignment into retrieval of paper authors and cited authors, and its experiments show scholarly dense embeddings narrowly beating lexical and general-purpose retrievers.

  13. A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools

    cs.LG 2025-06 unverdicted novelty 4.0 of 10

    This survey organizes foundation models, LLM agents, datasets, and tools in materials science into six task areas.

  14. LGAI-EMBEDDING-Preview Technical Report

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A Mistral-7B embedding model trained with in-context instructions, soft labels from an in-house retrieval pipeline, and margin-based hard-negative mining reports top-tier MTEB English v2 scores.

  15. A New Query Expansion Approach via Agent-Mediated Dialogic Inquiry

    cs.IR 2025-02 conditional novelty 4.0 of 10

    AMD uses three LLM agents (Socratic questioning, dialogic answering, reflective feedback) to generate and refine pseudo-answers for query expansion, reporting gains over prior methods on BEIR and TREC benchmarks.

Pith tools