Pith. sign in

REVIEW 13 cited by

Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.08037 v2 pith:Z5KZHLDG submitted 2022-12-15 cs.CL

classification cs.CL
keywords attributedllmsattributiondevelopmentevaluationlanguagelargemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have shown impressive results while requiring little or no direct supervision. Further, there is mounting evidence that LLMs may have potential in information-seeking scenarios. We believe the ability of an LLM to attribute the text that it generates is likely to be crucial in this setting. We formulate and study Attributed QA as a key first step in the development of attributed LLMs. We propose a reproducible evaluation framework for the task and benchmark a broad set of architectures. We take human annotations as a gold standard and show that a correlated automatic metric is suitable for development. Our experimental work gives concrete answers to two key questions (How to measure attribution?, and How well do current state-of-the-art methods perform on attribution?), and give some hints as to how to address a third (How to build LLMs with attribution?).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 26 citations worldwide. Full citation record

  1. ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    ProvenanceGuard detects when a claim in an MCP-based agent answer is supported somewhere but attributed to the wrong source, with block F1 0.802 and perfect detection on 50 controlled swaps.

  2. Ceci n'est pas une pipe: AI systems as semantic abstractions

    cs.AI 2026-07 conditional novelty 6.0 of 10

    AI systems are formalized as semantic abstractions whose claims are reliable only when supported by universal knowledge, source-derived knowledge, current effective knowledge, and explicit authority.

  3. CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion

    cs.LG 2026-07 accept novelty 6.0 of 10

    CollabEval turns model evaluation into low-rank matrix completion and uses the imputations as control variates to cut CI width and MSE at fixed annotation budget while preserving unbiasedness.

  4. Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

    cs.CL 2025-10 conditional novelty 6.0 of 10

    SLIM separates search and browse tools and summarizes trajectories every 50 turns, beating several open-source agentic search systems on BrowseComp and HLE with fewer tool calls and lower cost.

  5. Cost-Optimal Active AI Model Evaluation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    The paper derives budget-optimal rules for mixing cheap weak raters and expensive strong raters so that the mean rating is estimated with minimum variance, and shows large cost savings when example difficulty varies.

  6. Understanding Mental Models of Generative Conversational Search and The Effect of Interface Transparency

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Users of generative conversational search mostly hold abstract, incomplete mental models, and added interface transparency did not reliably improve those models or satisfaction.

  7. LAQuer: Localized Attribution Queries in Content-grounded Generation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LAQuer defines user-initiated, span-level attribution for grounded generation and shows it can cut the text users must read to verify a claim by about two orders of magnitude, at the cost of lower attribution accuracy.

  8. How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new GraphRAG evaluation framework using graph-grounded questions and bias-correction yields much smaller win rates than earlier reports, casting doubt on reported GraphRAG gains.

  9. MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LLM relevance judgments align with physician trainees on only about 45 to 66 percent of sentences, and pruning contexts to physician-labeled relevant sentences improves LLM and trainee accuracy.

  10. CAGE: Cognitive Attribution Graphs for Faithful Inline Citation Generation in Long-Form Question Answering

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Explicit cognitive attribution graphs before generation contract claim–document assignment space and yield SOTA faithful inline citations on long-form QA benchmarks.

  11. MedCite: Can Language Models Generate Verifiable Text for Medicine?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A two-pass combination of retrieval-augmented generation and post-hoc citation seeking improves citation precision and recall for medical question answering, and LLM-based attribution judges agree with physicians only...

  12. Pretrained LLMs Learn Multiple Types of Uncertainty

    cs.CL 2025-05 conditional novelty 5.0 of 10

    LLMs encode multiple dataset-specific linear directions in their hidden states that predict their own answer correctness, and these directions are nearly independent across benchmarks.

  13. REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing

    cs.CV 2025-05 conditional novelty 5.0 of 10

    REGen generates documentary teasers by fine-tuning an LLM to write a script with <QUOTE> markers, then a trained retriever fills each marker with the most relevant clip from the source video.

Pith tools