Pith. sign in

Paper Citation Record · LEDGER

CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2402.14809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.14809 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:50:12.337684Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 68ea7ff4-bd91-4c7b-ac97-382d98d49848 · inbound

Improving Video Generation with Human Feedback cites this paper.

Improving Video Generation with Human Feedback CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:30:02.737394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T15:30:02.578430Z digest=sha256:d49b2868ae5537c85fb85f9ba63fe910bd4602d2012f87a60230e6b99d9cd73b

Observation 6acbb55f-3aed-4f68-b65f-083e46dd1ea8 · inbound

LLMs can be easily Confused by Instructional Distractions cites this paper.

LLMs can be easily Confused by Instructional Distractions CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T10:50:12.337684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:50:12.337684Z digest=sha256:865fe6ada84ebe7a1986b5bca5f72507c1c4af54f28493a8e78768efa9cf7c91

Observation 7a419c50-c06a-4eac-9b93-1900ea188b43 · inbound

Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving cites this paper.

Step-Wise Formal Verification for LLM-Based Mathematical Problem Solving CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:48:56.468279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:48:56.468279Z digest=sha256:78182c2bea3374efc6c87996c7a207d1db7744e77ea8d758aa126fae142ad3b2

Observation 1bca06c8-7f32-4bb1-a072-615a8862484e · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.201852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:a7fa3e311cd5054e9a8a138e912e996e1ad643a0937b2cdbf2c55138035fe825

Observation 7bf96a30-5c1b-49ba-9e5c-4cc6c5e0f159 · inbound

VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation cites this paper.

VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:33.813200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:33.813200Z digest=sha256:123c398a1df80e3970096b9054790bbb27202803c73a224bccc5c6ed45f208ae

Observation 9a806e1b-7c91-41a8-8c31-fd9bfdcef0b0 · inbound

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization cites this paper.

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:16.816999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:14:16.816999Z digest=sha256:01b4d9b886177372774f69fd39e2a753a4edd471c0d32736821a29515af74351

Observation 6d143611-5771-45c2-b162-1a9dcf1eea7c · inbound

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback cites this paper.

RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:47:31.805450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:47:31.805450Z digest=sha256:91c496cd4d5466696c9721147e88e89ed9a9186aa216413555aaebcf5b962480

Observation daf54d4c-c5c4-4190-b04a-cbd843e76f0b · inbound

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability cites this paper.

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T01:00:27.481589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:00:27.481589Z digest=sha256:7af55911516f26eef6b541e47b953eb6697580ce7cc431882ef67e9f536ac22a

Observation 2086b781-cc8b-44aa-9d5e-0e6140c01720 · inbound

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents cites this paper.

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:07.916093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T03:21:30.732925Z digest=sha256:5ebb823f7a45b36ffbcd2abb39175906d6198b162ec138d99dd69123aa95b59d

Observation 2578d05a-2c2a-45e4-9621-83c45bd0c98d · inbound

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents cites this paper.

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents CriticBench: Benchmarking LLMs for Critique-Correct Reasoning

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:05:29.030156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T05:59:13.631078Z digest=sha256:5af940c93575cc07c5f9c405ea9c8d78e26e2b64b56774340fe5753e9cea94d1