REVIEW 15 cited by
SPECTER: Document-level Representation Learning using Citation-informed Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Representation learning is a critical ingredient for natural language processing systems. Recent Transformer language models like BERT learn powerful textual representations, but these models are targeted towards token- and sentence-level training objectives and do not leverage information on inter-document relatedness, which limits their document-level representation power. For applications on scientific documents, such as classification and recommendation, the embeddings power strong performance on end tasks. We propose SPECTER, a new method to generate document-level embedding of scientific documents based on pretraining a Transformer language model on a powerful signal of document-level relatedness: the citation graph. Unlike existing pretrained language models, SPECTER can be easily applied to downstream applications without task-specific fine-tuning. Additionally, to encourage further research on document-level models, we introduce SciDocs, a new evaluation benchmark consisting of seven document-level tasks ranging from citation prediction, to document classification and recommendation. We show that SPECTER outperforms a variety of competitive baselines on the benchmark.
Forward citations
Cited by 15 Pith papers
-
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
Conference accept/reject outcomes yield 15 operational ideation patterns that, as an LLM skill suite, improve automated-judged research-proposal quality over no-skill and generic-skill baselines.
-
Mutual Linearity in and out of Stationarity for Markov Jump Processes: A Trajectory-Based Approach
Trajectory-level linear response yields mutual linearity of observables under single-edge rate perturbation for Markov jump processes, including non-stationary state and counting observables.
-
BitNet Text Embeddings
BITEMBED trains 1.58-bit ternary-weight LLM embedders with contrastive pre-training, supervised distillation, and multi-precision output training, matching FP16 teachers within ~0.6 MMTEB points at ~2x CPU speed.
-
Col-Bandit: Query-Time Top-$K$ Estimation for Late-Interaction Retrieval
An adaptive confidence-bound cell-pruning method recovers the exhaustive MaxSim top-K with roughly one-quarter to one-third of the compute on BEIR and REAL-MM-RAG, at the price of a calibrated rather than certified de...
-
From scratch to silver: Creating trustworthy training data for patent-SDG classification using Large Language Models
A weak-supervision pipeline using LLM-extracted concepts and rank fusion creates silver-standard patent-to-SDG labels that recover known citation-derived associations and show high network modularity.
-
SPAR: Scholar Paper Retrieval with LLM-based Agents for Enhanced Academic Search
An agentic academic search system using LLM agents for query understanding, citation-chain exploration, and reranking outperforms prior retrieval baselines on two benchmarks.
-
Four Shades of Life Sciences: A Dataset for Disinformation Detection in the Life Sciences
Introduces FSoLS, a four-class labeled corpus of 2,603 full-text life-science articles, and benchmarks language models that classify disinformative texts with up to 98% F1.
-
MIR: Methodology Inspiration Retrieval for Scientific Research Problems
A new dataset and MAG-guided triplet-loss fine-tuning, plus LLM reranking, improves retrieval of methodologically inspirational papers for research proposals by several points over strong baselines.
-
Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs
An LLM-based pipeline extracts technology triples from arXiv and patent full text, tracks rising topic co-occurrence, and labels retrieval-augmented generation and conversational agents as emerging transformative tech...
-
A Comparative Study of Specialized LLMs as Dense Retrievers
Specialized Qwen2.5 7B models differ in dense retrieval quality: math and long-reasoning variants degrade performance, while coder and vision-language variants improve zero-shot text and code retrieval.
-
Extracting Information About Publication Venues Using Citation-Informed Transformers
SPECTER paper embeddings show that some CS venues are nearly indistinguishable and that several venue pairs converged between 2015 and 2023.
-
exHarmony: Authorship and Citations for Benchmarking the Reviewer Assignment Problem
exHarmony is a public benchmark that turns reviewer assignment into retrieval of paper authors and cited authors, and its experiments show scholarly dense embeddings narrowly beating lexical and general-purpose retrievers.
-
A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools
This survey organizes foundation models, LLM agents, datasets, and tools in materials science into six task areas.
-
LGAI-EMBEDDING-Preview Technical Report
A Mistral-7B embedding model trained with in-context instructions, soft labels from an in-house retrieval pipeline, and margin-based hard-negative mining reports top-tier MTEB English v2 scores.
-
A New Query Expansion Approach via Agent-Mediated Dialogic Inquiry
AMD uses three LLM agents (Socratic questioning, dialogic answering, reflective feedback) to generate and refine pseudo-answers for query expansion, reporting gains over prior methods on BEIR and TREC benchmarks.
Discussion (0). Continue with ORCID to comment.