Pith. sign in

Paper Citation Record · LEDGER

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework

As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2507.08459.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08459 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:24:25.365824Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9f3ca66-a025-414b-b31d-d8abe431ff86 · outbound

This paper cites GPT-4 Technical Report.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:25.964780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:25.964780Z digest=sha256:1eca9cf343f82d8caccd1803b8b075b544b7b609f7bfa2a8f394f9ef6426db44

Observation bceabd59-e50c-4f18-9e7f-45e9b482c200 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:52.691946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:26.044365Z digest=sha256:a97bfe0e3b7b00a164933f81c0d86302c8b064f107f4a0c3539ffd14b4c4f904

Observation 087caa4c-7d4a-44f4-a55c-aa2fdb1bbc54 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:26.134301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:26.134301Z digest=sha256:e7ed64c685625d0ca871dbea62df9ffbe5e0d30189afea9a82169aa7c4be9ea7

Observation daeef872-b5e5-4936-a120-bb95b156681a · outbound

This paper cites Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:26.267312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:26.267312Z digest=sha256:8f4b961f747bb939463a85f373ccb2bb85b5f900d16f3f19ec16650e677d62ab

Observation dd0c1e34-8861-4974-b231-e7639806c440 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:52.555856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:26.546769Z digest=sha256:4bb524cf193777d82a98f499b7e96d81d90be8222c1ff8f346e93302c31728e2

Observation ca89c17d-2ea3-4b50-a655-d9ba32c648f0 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:27.817012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:27.817012Z digest=sha256:c33b1a4895e616f25cb091bf9d9b24b4031546bc3008cc09bd2939a09beb2abd

Observation 873abe80-539f-48fc-8894-b77aa8dcbd3b · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:52.391359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:28.921162Z digest=sha256:61ec425ef933175080f6905fceb6a02a56bbf93d8391c14d0bf4a117d4a8c11e

Observation 4e139474-0b9b-416a-90d1-69540bcde926 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:29.038244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:29.038244Z digest=sha256:f7f9cb0a801db786692ba6af7de35bfcae63f2cc3fafbc6fc5b3a520868374c2

Observation 562f96b2-321b-45c1-b913-0a0329011e3a · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:52.113547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:29.156199Z digest=sha256:0e827c2dfa80d274a115167cc0619dae1be69f1322bff3fa911a7d3d94ca80fd

Observation 48bbf629-5b02-4954-b96c-268df6e33240 · outbound

This paper cites Evaluating LLMs at Detecting Errors in LLM Responses.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Evaluating LLMs at Detecting Errors in LLM Responses

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:29.325939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:29.325939Z digest=sha256:4f3951c80a9bccfc3fd664d62c1274c1ffb9c1f6faeee4a439206396c0c94a7b

Observation 57a8272d-7e70-4836-b9c1-cf6b5f51b41c · outbound

This paper cites When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:29.536585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:29.536585Z digest=sha256:d17110c0298e94bdc93dd9232bc442593e8c660958048ea39e8021d156cf7d06

Observation 7d6f9917-e9a8-4f87-9421-077fbc91221b · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:51.925597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:29.705529Z digest=sha256:3029d002c177cebcb48e641a36c120069715ab83b1c90801fd3033a0f209dc2b

Observation 5571812b-aa64-4fc1-b3b2-dcbcde42b081 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:29.831673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:29.831673Z digest=sha256:6d2b459bf43b5fd42e9b3264124852f8901a31076c821b85d87b668dc88891b5

Observation ba518ebf-4f5b-45a9-98e3-d4b01ab32bd2 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Adam: A Method for Stochastic Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:29.935279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:29.935279Z digest=sha256:1acc84b99a04f127bc77040c23e43341d4203cf35bd43c7469d714f33ac7bb82

Observation 2f88b4fc-64f0-4dcb-8586-2d357059a6bc · outbound

This paper cites Hurdles to Progress in Long-form Question Answering.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Hurdles to Progress in Long-form Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:30.057653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:30.057653Z digest=sha256:10a85e7a9f05eed2a3a9c9ee96af33e92fbb9837247f08f6caa6c0195a4197bc

Observation fa3ef39b-e07e-4236-9eb8-7daeb66490a9 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:30.169471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:30.169471Z digest=sha256:9de5033329006405e491cb92c33d08b674f0d550c3dc91cf681d2d7b0dbce8cd

Observation 735d5c09-e9e6-4558-8dcf-8a0435284eef · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:30.235940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:30.235940Z digest=sha256:bcef88ae20990ddc80317d6d1a4d71a9a5e0c98e206243d1db1c39ec837a2349

Observation be05120e-00ae-43bf-9769-bcbd31e2cae9 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:51.824361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:30.359679Z digest=sha256:1f629abc7fce210ca8555e653d3a96f663c3b0e5190c4b52513423eb8d55fe95

Observation 2a47dbe2-5407-4ed4-b574-e7ec36c316b5 · outbound

This paper cites AlignBench: Benchmarking Chinese Alignment of Large Language Models.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:30.510620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:30.510620Z digest=sha256:9d9db50e511b26ee8b5fae2fe01a02889caf54e0ab36411df5453130e7d607af

Observation c4f9a607-a61e-4c25-9b0b-7c0727169d30 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:51.710663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:30.634447Z digest=sha256:815bc92450b853733585ada34f8616ba09226663b3e62c7b643fab6efddf9535

Observation 5e30c80b-8e26-43c0-84ea-abf8d040141e · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:30.789112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:30.789112Z digest=sha256:bf4bf3adddd8ce620ef27dc1318244c84cc81109836a28c7f703d2d31a447b84

Observation 3f201c77-f884-4c8b-bfec-3a2fa0a0d2d9 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:51.591337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:30.924621Z digest=sha256:022f3b7101150496940c29355eb0187bb0c4211eeb1018d421d0a83afa23387f

Observation 80daa32e-9cd3-4174-ac5b-a29a1739c9a7 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:31.040317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:31.040317Z digest=sha256:6a7fc1ae6017172999b94942b3858f77d61906fafec8ff955e71900c6abebbbd

Observation eba66bd4-0b87-4b38-ba5d-34b3ba918326 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:31.175323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:31.175323Z digest=sha256:9268c6a92f14129ac9f9719fd2a9a92971420e66b44d6e906d1ee01dfe43256d

Observation 6274dd55-b188-4be7-bfbd-cf141b850aa3 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:51.488387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:31.298506Z digest=sha256:2cbc519de4fce1baa52de478df4961ad7ed6a6f98e2b1b6a7282e6817c5a9682

Observation 904dc09d-dc34-4f07-b765-b10a73164ad0 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:31.410452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:31.410452Z digest=sha256:9b47d4238b2bcb9d641d32f5fb1886d649abaa6134ec4b2db9e4197df86eae1e

Observation 07c841b8-0bba-451b-8d83-2bd5866d92ff · outbound

This paper cites Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:31.499480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:31.499480Z digest=sha256:ec578c487058c4c043e9d7848c5a7d8d67f077d22472fe663791e4038eb08644

Observation db74c65a-fc59-4d18-9e1c-b82356c29856 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:31.587962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:31.587962Z digest=sha256:a2388e764d08d990a8749dcbb4d61bb94724aabd156439240d998d6eb40e14c8

Observation dda5ec9b-2c13-4531-800e-83d248164164 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:51.396815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:31.705758Z digest=sha256:2669407fdf2a328b5416b892e3727592a18f69d0e4418c54133ef9009afd9d6d

Observation cf087154-cc14-41c2-9010-7e91381fee32 · outbound

This paper cites BLEURT: Learning Robust Metrics for Text Generation.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework BLEURT: Learning Robust Metrics for Text Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:31.866770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:31.866770Z digest=sha256:8317d1927a88f2f54dc513d61dc4a318e54b2b0710921828db2a1bf7006f7348

Observation 2bf083f1-b564-47bc-a6ea-d2d2109194a6 · outbound

This paper cites An Evaluation of Estimative Uncertainty in Large Language Models.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework An Evaluation of Estimative Uncertainty in Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:32.007106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:32.007106Z digest=sha256:5a539a5f6a5f6b79fcc5d30e7aca98643b019ac97d8d361c48b61aa027bceef5

Observation fc21d2d3-d84d-4142-ae6a-3c2f76315369 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:51.251462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:32.186651Z digest=sha256:bb0e2c557d6591f7f58a39a872dea8416d140b9a4d377fe2d5ec05d5bbddd360

Observation 9552394f-21cb-4785-8ec1-db3052d5d66e · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:51.095486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:32.283977Z digest=sha256:797eb5feacad729410143ce072fa57ea2d660ffefa1cd48a4138d7c3c41bd481

Observation af2b1c80-2ff6-43b1-b8da-20d7b264fe99 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:32.421179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:32.421179Z digest=sha256:2d7e6e646764ccfd7481ac3118e3f6b7007a31ee30913444341986928e4aabd6

Observation b9047dcf-9bc5-460a-b1de-1f413897c989 · outbound

This paper cites TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMs.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework TencentLLMEval: A Hierarchical Evaluation of Real-World Capabilities for Human-Aligned LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:32.541837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:32.541837Z digest=sha256:3b8f5bfba476330081a48857f0880a3e98b875fbc3935d51a0b0c78fb10ae34a

Observation 8e70f680-dd88-44ff-a16e-d9d201a75cf7 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Baichuan 2: Open Large-scale Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:32.668859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:32.668859Z digest=sha256:7b0044b2ba92df8a6f7387af03bd8c42f46872941eacf9ee50f5def2530bbe41

Observation 379b55f4-7887-46ec-9982-53634e19f60e · outbound

This paper cites Qwen2.5 Technical Report.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Qwen2.5 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:32.810642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:32.810642Z digest=sha256:217a8d7757bcb9cf932bdab2c1063676df3e2a3e16879631a9866ab7027f35c4

Observation d35a0f33-98a2-4ad1-b83f-8a89cc39c756 · outbound

This paper cites CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:24:26.000398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:32.952221Z digest=sha256:030da0ad9ccebca09aa10e86e6071a094ed3064977f28739d57274e2a15d0c05

Observation 591f0708-ece6-40b7-810c-6725e078873e · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:50.965748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:33.076234Z digest=sha256:5ab0afeca843c9d1cac4cf14e3a15f9a092d2476457b9f2b6a2785f2c4afff20

Observation 09dd2b4b-40be-4975-b7a2-157396507eef · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:50.779910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:33.161093Z digest=sha256:25d8484af312442c20c6023fc17a664834c82840e02ff24dc44de85b6cea72d1

Observation 7b20d8df-ac1b-496a-a097-1d89ee5bbfff · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:50.639692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:33.266794Z digest=sha256:81568e1f742b8b5296fccd1e2b8fdc4fe38c1ce04cf1dc093fb44ebc52813033

Observation a2954a49-8c1b-4732-9f4c-81e9249efe94 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:33.455203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:33.455203Z digest=sha256:d0ba69abc69f93d46d1c74c159b11ffbee198bb25845e45e57e5ef5ea27bd134

Observation f2b1e598-297e-45f7-8a1a-89b5a38ccf94 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:33.619260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:33.619260Z digest=sha256:ff3525691ab2bd011b08a44e36a0e7fd58d28435bb76c6e0eeb44393838acf4b

Observation b53d66ca-2f09-41b2-952a-1d404cc047e2 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:50.455163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:33.756105Z digest=sha256:45ffbc4667efc99cce865e3b21cfe00346ee809b31b96c7041a648ec8e9f41ee

Observation 934eadb4-79da-4b40-b25f-77842ad93286 · outbound

This paper cites How Language Model Hallucinations Can Snowball.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework How Language Model Hallucinations Can Snowball

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:33.820869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:33.820869Z digest=sha256:28585027de1bd175502ec2c98ac65c565c46d7de6cc33e7295b49ad6e67fe6a1

Observation 78f0b18e-bb5a-48c5-b3ef-75290c5973ba · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:26.580953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:33.928765Z digest=sha256:607e08ba3a44a17afbe16eee2cbe445f667fc5932ca5c556c0b80b31973a109e

Observation a2e2b92a-8308-4e67-8c34-04a422f9ed34 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework BERTScore: Evaluating Text Generation with BERT

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:34.057930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:34.057930Z digest=sha256:f69eee5251a8d91df359f9b3a8161167dd154064d21a95fe43981ad0c9056de0

Observation af0f8471-c3fc-4637-8ef5-efd3a5bbd331 · outbound

This paper cites an unresolved cited work.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:26.445899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:34.118699Z digest=sha256:4489346ec4e27c723dfbc19475de47dd881a1d21b5d0e404c5c45cbf1443cb59

Observation 35ee36bb-eb0c-467b-9e1b-542ec4faea67 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:34.253802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:34.253802Z digest=sha256:bde67f05a0ebd2fba67a5de2ae3ada80913ae1f2dadf3a08c1e2cef35948f504

Observation a80be7d0-c60d-499b-8818-86f40939115b · outbound

This paper cites TestNUC: Enhancing Test-Time Computing Approaches and Scaling through Neighboring Unlabeled Data Consistency.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework TestNUC: Enhancing Test-Time Computing Approaches and Scaling through Neighboring Unlabeled Data Consistency

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:24:25.603576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-06T18:23:34.988269Z digest=sha256:940f5013b05ea56d8362cfa641c03c817b8d24ba2e03d40e23f4d0c9b7bfdfdf

Observation da1fbe93-ca5d-4106-942e-38b30c5a861a · outbound

This paper cites LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:24.616624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:24.616624Z digest=sha256:19909b69a21226e96db9d40572163498715c5ede47e7568f154165731cf1342c

Observation 9e25c57e-e625-40ab-8761-8810550af946 · outbound

This paper cites online" 'onlinestring :=.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework online" 'onlinestring :=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:25.121221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:25.121221Z digest=sha256:cd786a55fa585d68112ed21f678a0ef38c178fa270a665d3caf293fa0c1315fa

Observation 56846acc-24c5-42e6-8a60-d09a7ba6dcec · outbound

This paper cites write newline.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:25.365824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:25.365824Z digest=sha256:122003f687c269a153300f550793397175ccb2a2890223fc54372368e97afa95

Pith citing papers

No inbound Pith citation observations are available.