REVIEW 13 cited by
Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have shown impressive results while requiring little or no direct supervision. Further, there is mounting evidence that LLMs may have potential in information-seeking scenarios. We believe the ability of an LLM to attribute the text that it generates is likely to be crucial in this setting. We formulate and study Attributed QA as a key first step in the development of attributed LLMs. We propose a reproducible evaluation framework for the task and benchmark a broad set of architectures. We take human annotations as a gold standard and show that a correlated automatic metric is suitable for development. Our experimental work gives concrete answers to two key questions (How to measure attribution?, and How well do current state-of-the-art methods perform on attribution?), and give some hints as to how to address a third (How to build LLMs with attribution?).
Forward citations
Cited by 13 Pith papers
-
ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
ProvenanceGuard detects when a claim in an MCP-based agent answer is supported somewhere but attributed to the wrong source, with block F1 0.802 and perfect detection on 50 controlled swaps.
-
Ceci n'est pas une pipe: AI systems as semantic abstractions
AI systems are formalized as semantic abstractions whose claims are reliable only when supported by universal knowledge, source-derived knowledge, current effective knowledge, and explicit authority.
-
CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion
CollabEval turns model evaluation into low-rank matrix completion and uses the imputations as control variates to cut CI width and MSE at fixed annotation budget while preserving unbiasedness.
-
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
SLIM separates search and browse tools and summarizes trajectories every 50 turns, beating several open-source agentic search systems on BrowseComp and HLE with fewer tool calls and lower cost.
-
Cost-Optimal Active AI Model Evaluation
The paper derives budget-optimal rules for mixing cheap weak raters and expensive strong raters so that the mean rating is estimated with minimum variance, and shows large cost savings when example difficulty varies.
-
Understanding Mental Models of Generative Conversational Search and The Effect of Interface Transparency
Users of generative conversational search mostly hold abstract, incomplete mental models, and added interface transparency did not reliably improve those models or satisfaction.
-
LAQuer: Localized Attribution Queries in Content-grounded Generation
LAQuer defines user-initiated, span-level attribution for grounded generation and shows it can cut the text users must read to verify a claim by about two orders of magnitude, at the cost of lower attribution accuracy.
-
How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG
A new GraphRAG evaluation framework using graph-grounded questions and bias-correction yields much smaller win rates than earlier reports, casting doubt on reported GraphRAG gains.
-
MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
LLM relevance judgments align with physician trainees on only about 45 to 66 percent of sentences, and pruning contexts to physician-labeled relevant sentences improves LLM and trainee accuracy.
-
CAGE: Cognitive Attribution Graphs for Faithful Inline Citation Generation in Long-Form Question Answering
Explicit cognitive attribution graphs before generation contract claim–document assignment space and yield SOTA faithful inline citations on long-form QA benchmarks.
-
MedCite: Can Language Models Generate Verifiable Text for Medicine?
A two-pass combination of retrieval-augmented generation and post-hoc citation seeking improves citation precision and recall for medical question answering, and LLM-based attribution judges agree with physicians only...
-
Pretrained LLMs Learn Multiple Types of Uncertainty
LLMs encode multiple dataset-specific linear directions in their hidden states that predict their own answer correctness, and these directions are nearly independent across benchmarks.
-
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
REGen generates documentary teasers by fine-tuning an LLM to write a script with <QUOTE> markers, then a trained retriever fills each marker with the most relevant clip from the source video.
Discussion (0). Continue with ORCID to comment.