Pith. sign in

Paper Citation Record · LEDGER

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives

As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2412.10220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10220 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:17:07.599252Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:25:46.925378Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:36:02.611276Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact3
  • verified fuzzy15
  • unresolved22
  • parse uncertain0
  • malformed identifier7
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66e8265a-312d-4d50-bcfb-756138f8f99a · outbound

This paper cites an unresolved cited work.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:17:08.108551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.458997Z digest=sha256:b3ee6e25f47f437b4a62fabea7f18292891ee22ecc64ece1d668d313dd17ba0d

Observation 9369f04f-3c22-4cf1-a56b-7868e9b8e011 · outbound

This paper cites an unresolved cited work.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:17:08.100788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.463646Z digest=sha256:d1a4735e8921e9f607afb636f367f9b8bf92c350457ba45c9a206215b3413159

Observation 100c07a3-073b-4e03-bfc9-2f220f0236f0 · outbound

This paper cites Flemish AI Research Program.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Flemish AI Research Program

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T16:17:07.953985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.467290Z digest=sha256:6b3d8191563194844897ceaf1657d835569841b9f99808bac7d31da26ae42fef

Observation cd76681c-7d66-48e2-87c5-d59b7ea23faf · outbound

This paper cites Lundberg and Su-In Lee.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Lundberg and Su-In Lee

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:08.093224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.471642Z digest=sha256:8435be4ce010d03816c2736c0f678558271b00d6d13abc449ea3aa788d341d60

Observation 066475d5-bf48-49a3-b03e-2192637f8faf · outbound

This paper cites ”why should i trust you?”: Explain- ing the predictions of any classifier.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives ”why should i trust you?”: Explain- ing the predictions of any classifier

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.475368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.475368Z digest=sha256:1a332b04782912b32ce66ac3eb8320ab00cbeb6566282f6eb16c10008d8911fb

Observation b73734ac-5de6-45a1-b424-00d5f4ae1e04 · outbound

This paper cites A value for n-person games.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives A value for n-person games

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.478794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.478794Z digest=sha256:e066216a3a2d267be58e6d9c9108a15f576db9019cd58e8bb6e5096577abdca8

Observation 9ced43c6-22ad-4380-9140-5644f9125d11 · outbound

This paper cites The inadequacy of shapley values for explainabil- ity, 2023.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives The inadequacy of shapley values for explainabil- ity, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:08.081653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.481445Z digest=sha256:17ba96d76180ef74ca4e6972601ce2134dd1e29b4dea10980c159a7ca02c7445

Observation 9ad79fb9-c267-417e-85d8-edbd5a8324eb · outbound

This paper cites Ex- plainability is not a game.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Ex- plainability is not a game

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.484973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.484973Z digest=sha256:64f1f9c8f5a9a9e2d32b897a1092657eebcdf8a84bc7b7d673408ced9d4fb7bc

Observation 49413a5f-6702-4cb8-91a9-8198cb441cc1 · outbound

This paper cites Natural language explanations for machine learning classification decisions.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Natural language explanations for machine learning classification decisions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.487993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.487993Z digest=sha256:30cbb7027f7cefaa6054162fcf0d1642c1ff1a3c892f7c975ccd62933f2cbcb3

Observation deceb318-f251-4882-a7e4-c8a7c4ff4d12 · outbound

This paper cites Tell Me a Story! Narrative-Driven XAI with Large Language Models.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Tell Me a Story! Narrative-Driven XAI with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.492051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.492051Z digest=sha256:8346f37db874f61de215756e6569a0a50471ad4ccc7dc2289c27a408e2c60a07

Observation 524364df-b89a-4dfc-ac09-9230b5223d68 · outbound

This paper cites LLMs for XAI: Future Directions for Explaining Explanations.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives LLMs for XAI: Future Directions for Explaining Explanations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.494861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.494861Z digest=sha256:9e7619f113f3cd660b040b756f4182a948a8f728d460e3003dcbdcb147a0c819

Observation 50f871d1-740f-4857-ad95-1f1bcdb6c489 · outbound

This paper cites Natural Language Counterfactual Explanations for Graphs Using Large Language Models.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Natural Language Counterfactual Explanations for Graphs Using Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-11T16:17:07.743925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.498907Z digest=sha256:47e7e6c8802d7c34396acf1300375acfea48f5f2e2a98f8aaa48ecde6b076c23

Observation 8add35a7-d315-4c68-a1c4-4c6cc4f2ab2d · outbound

This paper cites GraphNarrator: Generating Textual Explanations for Graph Neural Networks.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives GraphNarrator: Generating Textual Explanations for Graph Neural Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.502476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.502476Z digest=sha256:93ec2cdbece0a01804ae3d335f6105dbdbcc592d1310e6f34927589acc83f0ea

Observation bcc5ea74-d2d2-4da0-9243-819377d8c44b · outbound

This paper cites GraphXAIN: Narratives to Explain Graph Neural Networks.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives GraphXAIN: Narratives to Explain Graph Neural Networks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.506101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.506101Z digest=sha256:d0ff4cc7ef67339378e12fb42fe0c807a0b8020b3fa93c42e6848a75e3ff3166

Observation 8d0861c0-51d3-4439-ac24-10a72a6582af · outbound

This paper cites Faithful and plausible natu- ral language explanations for image classifi- cation: A pipeline approach.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Faithful and plausible natu- ral language explanations for image classifi- cation: A pipeline approach

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-11T16:17:07.508986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.508986Z digest=sha256:d74104ec3dc6f1a4d90e7110a6d5de69d32dc43622272c091f98df5cfca39e16

Observation f65c85bd-728f-4e63-a4f9-bae8820dd64e · outbound

This paper cites In-Context Explainers: Harnessing LLMs for Explaining Black Box Models.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives In-Context Explainers: Harnessing LLMs for Explaining Black Box Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.512632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.512632Z digest=sha256:95b0737c01c80890834063c1d73397dc92aec4d1fb24738fc5983c93fc26e688

Observation 00ae3619-fd37-4e18-8a74-c1197510ebd9 · outbound

This paper cites Explaining ma- chine learning models with interactive natural language conversations using talktomodel.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Explaining ma- chine learning models with interactive natural language conversations using talktomodel

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:08.074462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.516461Z digest=sha256:063be5b0c4067f45b94096bd076914b8fd9531bff2a1ed517862bbc1ef1c58d7

Observation 80d6ab0b-8b0c-4aef-875f-12ceea63da2d · outbound

This paper cites ME- TEOR: An automatic metric for MT evalu- ation with improved correlation with human judgments.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives ME- TEOR: An automatic metric for MT evalu- ation with improved correlation with human judgments

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:08.058600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.530835Z digest=sha256:f87740c88649fe543705751a23936c19fe3a3fd4060946fb9402814cb1978457

Observation 3460f186-b813-4886-88c3-85778c99a5fa · outbound

This paper cites Keane, Eoin M.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Keane, Eoin M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:08.066492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.521210Z digest=sha256:fc0bd12bf4d18f9f4f1a62636da7324e366f167320bfcd80debb821490b49f52

Observation 2fbded58-85d5-4a66-bd6a-bc8075889d01 · outbound

This paper cites Do models explain them- selves? Counterfactual simulatability of natural language explanations.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Do models explain them- selves? Counterfactual simulatability of natural language explanations

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:08.050942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.536547Z digest=sha256:5e90518a23edd36387ff339debb1e6eda9e1a6ee9cbe324b83b3c1172d828cf2

Observation 70fdefc5-6b79-40cd-ac59-42c48798b14c · outbound

This paper cites an unresolved cited work.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Unresolved cited work

Reference 21

Resolution
malformed identifier
no resolver link, observed 2026-08-11T16:17:07.525849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.525849Z digest=sha256:fe6690d200c77bde6aa2ef01f991b20c46f5b63ab9fcad40da4557222e31839f

Observation 88041425-3202-4a4e-b038-41cd32b37a42 · outbound

This paper cites BLEURT: Learning robust metrics for text generation.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives BLEURT: Learning robust metrics for text generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.528002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.528002Z digest=sha256:4188b7b56a881b51f518c783bad94319ed3c0088e883c67b5e783c413c53a128

Observation 15b8aa88-629c-4543-ae6c-bbf0700e4d07 · outbound

This paper cites The disagreement problem in ex- plainable machine learning: A practitioner’s perspective.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives The disagreement problem in ex- plainable machine learning: A practitioner’s perspective

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:08.033311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.550045Z digest=sha256:b00ea986bbdb91b7fab4b77c1b7d6923ee57716dc9da1ba657dbb88267106693

Observation 1403399b-4080-476a-9e3f-8fcbddee5ac0 · outbound

This paper cites F ActScore: Fine-grained atomic evaluation of factual precision in long form text generation.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives F ActScore: Fine-grained atomic evaluation of factual precision in long form text generation

Reference 24

Resolution
malformed identifier
no resolver link, observed 2026-08-11T16:17:07.533968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.533968Z digest=sha256:faa9c7ba26c158e54f3712ada3529df85a865bef726a213dd876a5e76a0c98a8

Observation 6c835f72-4275-4d2e-b9b4-426295af1f3a · outbound

This paper cites PRobELM: Plausibility ranking evaluation for language models.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives PRobELM: Plausibility ranking evaluation for language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:08.022375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.556439Z digest=sha256:c46bdec2a93f22a68bae9267819e820396bf8d25bcbe959d04d7674e4983a96a

Observation 24136dbe-fe04-4b8d-a2a1-a58fbf74c0d6 · outbound

This paper cites A Survey on Natural Language Counterfactual Generation.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives A Survey on Natural Language Counterfactual Generation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-11T16:17:07.693937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.558746Z digest=sha256:16852f8b14d12141b3fab35af36454b5332d24a9ae3d59d11eb59801171e8496

Observation 60bd9c9f-41c1-4e4d-898b-e54c2fdf6162 · outbound

This paper cites You Can Generate It Again: Data-to-Text Generation with Verification and Correction Prompting.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives You Can Generate It Again: Data-to-Text Generation with Verification and Correction Prompting

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-11T16:17:07.709487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.543352Z digest=sha256:729ff91afc97c52ec30bd4b05f5b2c8af6a4b6dfba68882a377914b0ab328178

Observation e249ad58-702f-42e6-a1f1-85fb671a1de3 · outbound

This paper cites The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.546973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.546973Z digest=sha256:9e22f1f05bea2ebe32c7e2841c4a18b9e0a5d39c0f873fbf1d0afffad1615491

Observation 2e1eb895-4274-40c0-b4b4-964cfe3e99eb · outbound

This paper cites Sentence- bert: Sentence embeddings using siamese bert-networks.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Sentence- bert: Sentence embeddings using siamese bert-networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:08.014238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.568624Z digest=sha256:51c044faec1d706ae282657d62094495ed03fd20a652b64445f854984e3050da

Observation d1972177-6ae1-4774-bb6c-5862a0a41aeb · outbound

This paper cites Towards few-shot fact-checking via perplexity.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Towards few-shot fact-checking via perplexity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.553636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.553636Z digest=sha256:6fb93ab28a9e55a3df0a81323d167bb7a844f4876371324cae55e4214bf14ce0

Observation 72b38990-3eef-4e12-81ad-7f6e730e9a4a · outbound

This paper cites GPT-4 Technical Report.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives GPT-4 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.576150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.576150Z digest=sha256:ce62c317f06d4ee3a69cd914bb25229fa005c955a7437641cc54931d077ab5a6

Observation 34835850-e086-40d9-aeb6-43a911d25f3a · outbound

This paper cites Claude sonnet 3.5.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Claude sonnet 3.5

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:07.999887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.580124Z digest=sha256:891a1d5083485a93d5c7673ba2aaddc14688579afd4b6de08a1c1ea17ace27d9

Observation df8d44f5-21e0-4902-95f6-bcf83eb50d43 · outbound

This paper cites Perplexity from PLM Is Unreliable for Evaluating Text Quality.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Perplexity from PLM Is Unreliable for Evaluating Text Quality

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.562462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.562462Z digest=sha256:cf314ca61f0fa186ac34ab18de0f1a1142793a809888c9ff181cb70015c243c0

Observation da3b3d94-ac13-40f7-8456-cc48b14adc3f · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Efficient Estimation of Word Representations in Vector Space

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.565772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.565772Z digest=sha256:a900bbfbb38b364731ec2416c508c5001faee5e62150e417730bb9c85ef0c5ae

Observation 592a9244-b580-4898-aba6-188aca8c656f · outbound

This paper cites Mistral large 2.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Mistral large 2

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:07.979102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.586417Z digest=sha256:e4967b08003a9e10a5fa3ad59975d6e9188163aa005eedae2ca44c39a4887b2f

Observation e04ac323-90c7-487e-94fc-275ee54e9ca6 · outbound

This paper cites Zhang, Mark Har- man, and Meng Wang.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Zhang, Mark Har- man, and Meng Wang

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.592213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.592213Z digest=sha256:4c0d7163b5d632a0c4d56b3793e9025780c9b0bbbe63bfa8175d08dd89399de7

Observation fe4b8619-b339-47a3-b923-d630d7b01e23 · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.573421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.573421Z digest=sha256:182fa9a37d5a98de2e7deb53848f680a78fc5778ce40aae4cf9ca8137d0c406b

Observation 48258627-7921-4712-bed7-593597329f0d · outbound

This paper cites Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:07.962454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.596922Z digest=sha256:e5f288817bde4ac75aae89d3db14566953cbdc2bf7198c821d7364a437433fbf

Observation e4af75fb-0b5a-4fe1-a1ed-d181f3b854b7 · outbound

This paper cites Context-faithful prompting for large language models.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Context-faithful prompting for large language models

Reference 39

Resolution
malformed identifier
no resolver link, observed 2026-08-11T16:17:07.599252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.599252Z digest=sha256:815e30f1aa1f053186485fd45c0e350e85c85d877636cdd369a92c2dea8e87f7

Observation 82bafdfe-e3be-47da-a95d-ef1fc45eebde · outbound

This paper cites Llama 3 model card.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Llama 3 model card

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.582600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.582600Z digest=sha256:40c567797932069117d44f057974c8cd02c770a72fc132985e329484fa87d17d

Observation 8c428be9-4abb-4c61-a14a-ae9af0b4199e · outbound

This paper cites The llama 3 herd of mod- els, 2024.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives The llama 3 herd of mod- els, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:07.987051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.584489Z digest=sha256:e915d5542cfe824130418923d59fb9bcbe259ae4999474931646401bbd55f0f7

Observation b975b56b-0f68-4a67-a197-23a7941a6e87 · outbound

This paper cites Entity-based knowledge con- flicts in question answering.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Entity-based knowledge con- flicts in question answering

Reference 45

Resolution
malformed identifier
no resolver link, observed 2026-08-11T16:17:07.594559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.594559Z digest=sha256:d732073657425de3f479fa44c9e7758d39da49b28bb9d325ba68555077a87465

Observation 532d0c77-1923-4217-964a-a52dad5963ae · outbound

This paper cites 19 org/CorpusID:201646309.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives 19 org/CorpusID:201646309

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:08.006627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.571095Z digest=sha256:0afeb7a5d6b82428b13a365d692f8bab647a75e339d2003188c2639b3f55ed94

Observation 0c52d410-b766-4a0c-bc5c-1417b54f10fa · outbound

This paper cites URL https: //doi.org/10.24963/ijcai.2021/609.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives URL https: //doi.org/10.24963/ijcai.2021/609

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T16:17:07.523664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.523664Z digest=sha256:d0a11e739b7c4932620980ebb2fc96b4b9e4e946623d6bbd7db45e3785428f11

Observation a7508e05-4ea2-42db-9c91-289addbcbf2d · outbound

This paper cites doi:10.1038/s42256- 023-00692-8.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives doi:10.1038/s42256- 023-00692-8

Reference 2023

Resolution
malformed identifier
no resolver link, observed 2026-08-11T16:17:07.518635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:17:07.518635Z digest=sha256:2f4009f78fc061f3dc3a4a614f6fffbbbfbf710abd50ed208bc18f492fb478f5

Observation 9a678ebe-0199-40d4-a588-da6812b0cc67 · outbound

This paper cites an unresolved cited work.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:17:08.041654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.539365Z digest=sha256:0ea1a42424de623cc8ac449133a153c083c8cb3a302b652540f951abdfc81a1a

Observation ca98e4fd-328e-43dc-bcd2-a8cfd1cdd7b3 · outbound

This paper cites URL https://huggingface.co/ mistralai/Mistral-Large-Instruct-2407.

How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives URL https://huggingface.co/ mistralai/Mistral-Large-Instruct-2407

Reference 2407

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:17:07.971441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T16:17:07.589377Z digest=sha256:dcd74c75a5a0bfda6ceaf2118e18563d179a8358d9455306b4893007c8f84465

Pith citing papers

Observation 3ecd87ac-c046-4939-811b-e477401c76f1 · inbound

A Two-Stage LLM Framework for Accessible and Verified XAI Explanations cites this paper.

A Two-Stage LLM Framework for Accessible and Verified XAI Explanations How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:36:02.614307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T15:25:46.925378Z digest=sha256:ee888c30d464e35bcefb638e84a736bfbd3b3a1ecd36e942927a28fe16511348

Observation 94a37fcc-18ec-4d54-b785-7b9f871ab5e1 · inbound

On the Importance and Evaluation of Narrativity in Natural Language AI Explanations cites this paper.

On the Importance and Evaluation of Narrativity in Natural Language AI Explanations How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.985155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T05:21:12.989059Z digest=sha256:962e0a3c3ab09d6add4e158e0884c8ff81a2cd70009e7382a81acc2a6c96efcb