Pith. sign in

Paper Citation Record · LEDGER

DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2309.17167.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.17167 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:51:15.504623Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:18:22.614907Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 80abcfa1-0db8-4723-ad31-ad77ce6bf351 · inbound

AI Alignment: A Comprehensive Survey cites this paper.

AI Alignment: A Comprehensive Survey DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:28:49.009396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T14:28:48.987140Z digest=sha256:da441a2ac7f85c3f4d731e61f14e40fcb9466aebf1dea472cc716f7bb0eff45d

Observation d096e8aa-526c-456e-9285-3ca85a935ca0 · inbound

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey cites this paper.

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:16:41.701975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T23:16:41.679855Z digest=sha256:daa3cc78c3caf4ff51172e4b40f0d1689852d841c94df9d3a9be6f7e27e801df

Observation 4d3f199f-97ef-4ba1-aebd-2a112b542616 · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.504623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.504623Z digest=sha256:e01e4a567fa924df1de4b6395255c97173d2f96c464472dadd9230f4b1ce6376

Observation ffb450d2-df90-4b18-9933-f04b013d0295 · inbound

Automated Capability Discovery via Foundation Model Self-Exploration cites this paper.

Automated Capability Discovery via Foundation Model Self-Exploration DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:55.420540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:55.420540Z digest=sha256:3b06b258baa7cb8366eb9e579b26c20ae5da2941538ca2ac4b1ad499f1fc80a9

Observation bcc4f788-8ab4-443c-b4bf-217c5acd20de · inbound

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation cites this paper.

Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.859281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.859281Z digest=sha256:2326ff508b18b282d129f4e4ff283a7fbfaac982eedfdfe3f3fd64ec5f6edc1c

Observation 733e74a4-d303-4d0a-9215-9bd1fdd42f1d · inbound

Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era cites this paper.

Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T15:43:12.862651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:43:12.862651Z digest=sha256:7e1b7ea9b51a93dfe328e694cd03c53c6f84fb03e112c8137d147290d619ea0f

Observation fd6e6606-853a-45cf-9a72-7e3e07d768a2 · inbound

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks cites this paper.

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:29.575749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T04:27:11.735657Z digest=sha256:b5abfedcacbe0fac9330ff079838f488edf2d6ce8b290e2c681ac358e0e44757

Observation d8b1cad9-4724-4f5e-a61c-630045c9aaf5 · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.745491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:575f418329868a96faaa191c78bf800947fd3323faa3a80400a8ab1091867ef6

Observation 72ba854a-499a-49dd-8308-cf390beb281c · inbound

From Text to DSL: Evaluating Grammar-Based Model Generation Using Open LLMs cites this paper.

From Text to DSL: Evaluating Grammar-Based Model Generation Using Open LLMs DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T16:43:34.739101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T16:40:08.788179Z digest=sha256:152f5ac2694ca4de795e85b36ca9047818a8d6925b793bfd9ac64a63643d4289

Observation d15d5841-55e1-4ab4-aac9-a2a998bfcdcb · inbound

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations cites this paper.

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:26:09.983958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T06:25:06.024542Z digest=sha256:342ba1f92e41df39c26ef15b777c11de43a54768bdbdab9a845c77ec2eae31d6

Observation 1e0e4761-68dd-421c-af32-243da1e6ddd5 · inbound

A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction cites this paper.

A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 43

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T14:18:22.616286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T14:08:50.462880Z digest=sha256:8cb412e2f32afcf2c09ca502e12ce2f25301664d7149084d32b8ee4e8f6aa2e0