Pith. sign in

Paper Citation Record · LEDGER

ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2212.07919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.07919 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:24.431078Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

28
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3f805292-2812-4c37-b4cd-0359cbb56397 · inbound

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark cites this paper.

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:24.431078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:24.431078Z digest=sha256:db41a961e54b714d3788aa7ed94381e51e09201cc300f2640534cdeb1fec0c33

Observation ee8e5cc5-8649-48d6-886a-a4b8f18fb5a6 · inbound

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects cites this paper.

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:04.297202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:04.297202Z digest=sha256:2615727b2ac5feea67f8dbe8131bffff5c02371e769dd49fa0ba9bbcd41f0ffa

Observation fe22c061-ad7b-4ba9-9e23-205a1610624f · inbound

Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains cites this paper.

Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:30.453495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:30.453495Z digest=sha256:1032622f0791570e5c50835acad280dd5c57453eb1f4b522215996683d9bcf80

Observation fb7d36f6-e76c-4daa-83bb-36b16ce378b4 · inbound

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation cites this paper.

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:08.212502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:51:08.212502Z digest=sha256:939248a478fddd476d66083b72333046c05d014e4bfd803421801d05aeea6a5a

Observation 04f788f9-ede6-409c-a8d0-e0e8d9e68fdf · inbound

VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation cites this paper.

VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:33.692878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:33.692878Z digest=sha256:c96632a00037d3ea9d2f879de885ed9784b4cec848306c45287796d9f48a239e

Observation 4a6dd9a8-749d-476d-b2fd-8ef1e599dc4a · inbound

Rethinking Human Preference Evaluation of LLM Rationales cites this paper.

Rethinking Human Preference Evaluation of LLM Rationales ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T17:15:27.868447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:15:27.868447Z digest=sha256:bb5cf193cafbf11f4d3eb191a6b5c55d78a6e38d7da10bbe6f962e8e1e0a84f2

Observation d0a98575-162f-4037-b710-4260fb60d90d · inbound

Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations cites this paper.

Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:23.684248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T12:58:50.538826Z digest=sha256:62e64935d20a6a791a8a80ccc58d8bc79c2db66d042841ce56d1f398230d737b

Observation 6a572793-a0ee-4424-ac17-5b8c3f6d972a · inbound

Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring cites this paper.

Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T16:14:28.968859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:14:28.968859Z digest=sha256:a5ef235d29afdde74a1e321e7c152ba9aa16db17c7715ef0d30e9afd8271e9e7

Observation 1b783235-d6ef-41dc-b722-ff326ce7efa8 · inbound

Diagnosing Pathological Chain-of-Thought in Reasoning Models cites this paper.

Diagnosing Pathological Chain-of-Thought in Reasoning Models ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T23:26:29.354201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:26:29.354201Z digest=sha256:8e8ce6dbfdf00ea114db792746ab5db65280f43f72673d0dace07f35afe1a00f

Observation 4278645e-906e-45e5-8d16-3be97360b09b · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:47:15.328082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:e792ea81b64a72fed1f9f13c1445cb5349389dd44e7815a10536ce0317ace194

Observation 45b69f89-53be-4c34-95c8-8384bed91f85 · inbound

Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework cites this paper.

Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:55:44.645241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:53:12.473010Z digest=sha256:f0b9571f8a7ec3b561480b08bb9f01c2312c03f5a0059c531252001cfaebaec0

Observation e92b42ec-b7e8-40d5-9e4b-ac03b7a92cd9 · inbound

From Hallucination to Structure Snowballing: The Alignment Tax of Constrained Decoding in LLM Reflection cites this paper.

From Hallucination to Structure Snowballing: The Alignment Tax of Constrained Decoding in LLM Reflection ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:30:51.495749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:04:32.931596Z digest=sha256:821b0d76abe291a74437d24f20256dc85fdd2588dfb601797a13fbe389a218e4

Observation 91f99d7a-2a47-4957-b787-2caae3e80e74 · inbound

Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction cites this paper.

Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:37:36.780083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:48:39.832057Z digest=sha256:32109f2bcf130de9fd2827a8bf9b1a326a7eb1dd139150a37a0a0caebaa98814

Observation cd110505-f34c-4fbf-b9db-7a0b6c0de025 · inbound

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling cites this paper.

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:44:22.077461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T07:37:38.371341Z digest=sha256:f28adb5aacd231123fbfb5732deece5dea16054b8166d6ee33f60ddb2ee9b2d3

Observation e118551f-6c45-4c67-9cc5-8e24d7e4d013 · inbound

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning cites this paper.

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:25:41.607755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T05:33:52.027771Z digest=sha256:cd1515870f07b35a9b052c94c9330a48afa2e9791506f2f7512b186267c64fff

Observation cb15932c-4aba-4345-b557-31fa45f113f4 · inbound

More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness cites this paper.

More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T01:00:54.813255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:00:54.813255Z digest=sha256:eb3f10dd1a7bd83765bbfc7879c03ecbba26144f5fa8b1a95ef14ab7ad784b2c