Pith. sign in

Paper Citation Record · LEDGER

Dynamic benchmarking framework for LLM-based conversational data capture

As of 14 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2502.04349.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04349 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:15:05.313014Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:43:16.241588Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T14:43:16.748149Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3d79fd6-07d1-4ed1-ba0a-6dc6fd0c52f0 · outbound

This paper cites Beyond Prompts: Dynamic Conversational Benchmarking of Large Language Models.

Dynamic benchmarking framework for LLM-based conversational data capture Beyond Prompts: Dynamic Conversational Benchmarking of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.232557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.232557Z digest=sha256:3581265ea323799a7cdef44bb53596b9960136361c268fdabfca54b9d8c6f6d1

Observation 46aa19f4-babb-48ad-8329-088b7f2e7d6a · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

Dynamic benchmarking framework for LLM-based conversational data capture Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.244587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.244587Z digest=sha256:65b8bb11db0037309b79980525d310f5fb4d33a051241037f66b75eecc74e443

Observation 0de8e0f8-411f-44ab-8d66-ee52e1785173 · outbound

This paper cites Making Pre-trained Language Models Better Few-shot Learners.

Dynamic benchmarking framework for LLM-based conversational data capture Making Pre-trained Language Models Better Few-shot Learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.255887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.255887Z digest=sha256:c7d6574e2d64084322146af4c7bcdb75c810428bb96e9c481f1d59cb11b80994

Observation 44c35fb5-b18c-4166-9b13-e5f582701fdd · outbound

This paper cites an unresolved cited work.

Dynamic benchmarking framework for LLM-based conversational data capture Unresolved cited work

Reference 10

Resolution
verified exact
doi, observed 2026-08-09T12:15:05.446056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-09T12:15:05.259451Z digest=sha256:26cc5fd12da9d4c6b6409a85a943000735ce01cb129b02c2ee457ebf1dc4dad3

Observation b346bbd0-3f7f-4506-adca-1e69265dabf0 · outbound

This paper cites an unresolved cited work.

Dynamic benchmarking framework for LLM-based conversational data capture Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:15:05.622562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-09T12:15:05.262498Z digest=sha256:25dd0778a73d03eafbf3fbf8ad37925a264b14dba0421f1ef86ed13d41731d01

Observation 40472905-a9fa-40a1-86a5-d8872f4790a1 · outbound

This paper cites Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions.

Dynamic benchmarking framework for LLM-based conversational data capture Unveiling LLM Evaluation Focused on Metrics: Challenges and Solutions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.270152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.270152Z digest=sha256:d9e62cfd4c076062decfbd767dbfc5b2dfc2b2f574f62ff9755a3ee325bc737c

Observation 21c9d41a-c017-4551-9245-53e35e6eb829 · outbound

This paper cites Challenges and Applications of Large Language Models.

Dynamic benchmarking framework for LLM-based conversational data capture Challenges and Applications of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.274234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.274234Z digest=sha256:818400eb66a6166d6aae513bc7f23b2ebbe391c6e62099e23f0d407271c02b69

Observation 289baa1b-57dd-4d0f-b70a-abf23f01faa7 · outbound

This paper cites In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.).

Dynamic benchmarking framework for LLM-based conversational data capture In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.)

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.278228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.278228Z digest=sha256:f5e8863ff4334e3889d69c152fa971b8fb081f80ec118ce9c45711759f8e216c

Observation 59bd7b06-6dc2-40d8-b27a-b60826f6dcdf · outbound

This paper cites Mind the Gap: Assessing Temporal Generalization in Neural Language Models.

Dynamic benchmarking framework for LLM-based conversational data capture Mind the Gap: Assessing Temporal Generalization in Neural Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.281624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.281624Z digest=sha256:2a4f57b8e0ba63e8636f93526544342579a449fc9242077c987ffdd79f677854

Observation f84393a5-0c02-40ba-9418-dedb87497952 · outbound

This paper cites an unresolved cited work.

Dynamic benchmarking framework for LLM-based conversational data capture Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:15:05.607361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-09T12:15:05.285508Z digest=sha256:0b1b310900413e6669a71b864ee132b314d2bb660fb7dfaff2f4487d10872ae5

Observation a9d1dcf6-68e8-4fb4-9123-3b82a2aeb128 · outbound

This paper cites GPT-4o System Card.

Dynamic benchmarking framework for LLM-based conversational data capture GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.289325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.289325Z digest=sha256:cd2436aafa12197a8d027abf8b84705f9114da9a62fa1b9c6add557533b486d2

Observation 05360c84-1bbe-48b5-80f3-edc63783cf52 · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

Dynamic benchmarking framework for LLM-based conversational data capture Generative Agents: Interactive Simulacra of Human Behavior

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.297478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.297478Z digest=sha256:c17d6eae46aae4b299382217ad9438cb17da94959a67122290f63d8bac043d00

Observation 1e623876-1f18-41f0-a666-c5a59816843e · outbound

This paper cites CoQA: A Conversational Question Answering Challenge.

Dynamic benchmarking framework for LLM-based conversational data capture CoQA: A Conversational Question Answering Challenge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.301435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.301435Z digest=sha256:04aac79d630893a7435eec4853f0187173bd1e67df7cf6ea5c1519f2c2eaa464

Observation 6be2bb4f-c70a-4819-a5ac-86be86f64291 · outbound

This paper cites Measuring Semantic Coherence of a Conversation.

Dynamic benchmarking framework for LLM-based conversational data capture Measuring Semantic Coherence of a Conversation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.305901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.305901Z digest=sha256:807e63154f637b4178ec8ab7253690489641327bda39f0ae4bcb3d97248206cd

Observation 492fdea4-bdbb-4a8f-ae12-473b1ae492b1 · outbound

This paper cites an unresolved cited work.

Dynamic benchmarking framework for LLM-based conversational data capture Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:15:05.596053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-09T12:15:05.309715Z digest=sha256:0da69cf0a5e0f714a1fa5a9935af9319876b69b09726313991fa5d1f773adbf0

Observation e1d9affa-e025-4fc4-b69e-e0e77831ee06 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

Dynamic benchmarking framework for LLM-based conversational data capture MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.313014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.313014Z digest=sha256:86193e8af2e198eb54bbcfbe5f94d445b062bc2e0fdf2f4f32fda9559197e0db

Observation c25b8169-3d8b-46b9-8fb2-a2c4448a6d66 · outbound

This paper cites O’Brien, Carrie J.

Dynamic benchmarking framework for LLM-based conversational data capture O’Brien, Carrie J

Reference 311

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T12:15:05.584572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-09T12:15:05.294056Z digest=sha256:4eb2f83d122343039eaf25ee68f34c80c4cf948814cb6f725dcddd8a196754fc

Observation 0e5577cb-cea9-4079-8fcc-20d986778bfb · outbound

This paper cites 2016), Dynamic benchmarking framework for LLM-based conversational data capture 463–471.

Dynamic benchmarking framework for LLM-based conversational data capture 2016), Dynamic benchmarking framework for LLM-based conversational data capture 463–471

Reference 2016

Resolution
verified exact
doi, observed 2026-08-09T12:15:05.434085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-09T12:15:05.265740Z digest=sha256:a44cf535b69e26d95cb87d4c7035f8ec74a53f0144c5483b0dd7d9fd2108193e

Observation 1adef826-c6b8-403a-b8b2-a6bd28e12e0e · outbound

This paper cites InDesign, User Experience, and Usability: The- ory and Practice, Aaron Marcus and Wentao Wang (Eds.).

Dynamic benchmarking framework for LLM-based conversational data capture InDesign, User Experience, and Usability: The- ory and Practice, Aaron Marcus and Wentao Wang (Eds.)

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.248502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.248502Z digest=sha256:519453e5e7d049e4e348221732de3304747a0c86f26101d81ccb7b19bd49e730

Observation 2062522c-b1e0-44fc-89c4-0e1fe47183b9 · outbound

This paper cites InProceedings of the 2019 Conference of the North.

Dynamic benchmarking framework for LLM-based conversational data capture InProceedings of the 2019 Conference of the North

Reference 2019

Resolution
verified exact
doi, observed 2026-08-09T12:15:05.468244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-09T12:15:05.252276Z digest=sha256:9cb18c468291693afdcf55693554c9068197e82fc03959f6ef6d5fe116f64eef

Observation ecb8f8d7-c038-4436-8854-55dddb595b61 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Dynamic benchmarking framework for LLM-based conversational data capture Evaluating Large Language Models Trained on Code

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.240659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.240659Z digest=sha256:7a74d00ba2b314fee3a357ff03d26d3d18f94717a3e2f5740bc0127b13f1518c

Observation c7fd95b5-b416-45a6-8129-07d2b8adb15f · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Dynamic benchmarking framework for LLM-based conversational data capture On the Opportunities and Risks of Foundation Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.223946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.223946Z digest=sha256:6a0270aac1787b5ac1c88a292c2fb2c4ad2804479cb6fd2a4871d0e4595ba1f2

Observation 3ebbb4c2-02b4-4766-b7bb-fddb2c6c9ff2 · outbound

This paper cites Clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational Agents.

Dynamic benchmarking framework for LLM-based conversational data capture Clembench: Using Game Play to Evaluate Chat-Optimized Language Models as Conversational Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T12:15:05.236457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:15:05.236457Z digest=sha256:30405e2e6bdd746e04d8d8f3023923f4b71486581c22503ef99c78730802e56e

Observation cf4d25b3-a630-4b69-af9d-0d1c66112866 · outbound

This paper cites Large Language Model in Financial Regulatory Interpretation.

Dynamic benchmarking framework for LLM-based conversational data capture Large Language Model in Financial Regulatory Interpretation

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-09T12:15:05.533006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-09T12:15:05.228752Z digest=sha256:e36c785b676d4652a312758b41fddad3b93beb4281c2d10e81838232dc5bc355

Pith citing papers

Observation 46c991f9-9705-47d4-bc29-d02ccb290896 · inbound

ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents cites this paper.

ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents Dynamic benchmarking framework for LLM-based conversational data capture

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T14:43:16.752900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T14:43:16.241588Z digest=sha256:472055650b9e8ed69dc77549e7856ec483f02dacd88598b5f4f1fbd26776bc9d