Pith. sign in

Paper Citation Record · LEDGER

ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2212.07919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.07919 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:24.431078Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

28
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3f805292-2812-4c37-b4cd-0359cbb56397 · inbound

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark cites this paper.

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:24.431078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:24.431078Z digest=sha256:db41a961e54b714d3788aa7ed94381e51e09201cc300f2640534cdeb1fec0c33

Observation ee8e5cc5-8649-48d6-886a-a4b8f18fb5a6 · inbound

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects cites this paper.

Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:04.297202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:04.297202Z digest=sha256:2615727b2ac5feea67f8dbe8131bffff5c02371e769dd49fa0ba9bbcd41f0ffa

Observation fe22c061-ad7b-4ba9-9e23-205a1610624f · inbound

Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains cites this paper.

Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:30.453495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:30.453495Z digest=sha256:1032622f0791570e5c50835acad280dd5c57453eb1f4b522215996683d9bcf80

Observation fb7d36f6-e76c-4daa-83bb-36b16ce378b4 · inbound

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation cites this paper.

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:08.212502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:51:08.212502Z digest=sha256:4d9010bfe34ae5512af1e82475c276ae382a06c5e7c378a660184eb18ad83b7e

Observation 04f788f9-ede6-409c-a8d0-e0e8d9e68fdf · inbound

VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation cites this paper.

VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:33.692878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:33.692878Z digest=sha256:e205b6926f32afc6da2788f36be38b6867b3527d84c4432ce31440b63acd3f13

Observation 4a6dd9a8-749d-476d-b2fd-8ef1e599dc4a · inbound

Rethinking Human Preference Evaluation of LLM Rationales cites this paper.

Rethinking Human Preference Evaluation of LLM Rationales ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T17:15:27.868447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:15:27.868447Z digest=sha256:bb5cf193cafbf11f4d3eb191a6b5c55d78a6e38d7da10bbe6f962e8e1e0a84f2

Observation d0a98575-162f-4037-b710-4260fb60d90d · inbound

Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations cites this paper.

Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:23.684248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T12:58:50.538826Z digest=sha256:4ff07c795915c44575f80073e2ce23bb8b008461162961c25d7da22057fdf348

Observation 6a572793-a0ee-4424-ac17-5b8c3f6d972a · inbound

Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring cites this paper.

Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T16:14:28.968859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:14:28.968859Z digest=sha256:a5ef235d29afdde74a1e321e7c152ba9aa16db17c7715ef0d30e9afd8271e9e7

Observation 1b783235-d6ef-41dc-b722-ff326ce7efa8 · inbound

Diagnosing Pathological Chain-of-Thought in Reasoning Models cites this paper.

Diagnosing Pathological Chain-of-Thought in Reasoning Models ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T23:26:29.354201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:26:29.354201Z digest=sha256:8e8ce6dbfdf00ea114db792746ab5db65280f43f72673d0dace07f35afe1a00f

Observation 4278645e-906e-45e5-8d16-3be97360b09b · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:47:15.328082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:5c07188dacf6712ce1e1d8728074d98153c6eed59ca134952ffe14be7be7c487

Observation 45b69f89-53be-4c34-95c8-8384bed91f85 · inbound

Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework cites this paper.

Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T18:55:44.645241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:53:12.473010Z digest=sha256:d4c7622559f0982a3ce7f44e98d9f8be8234dffb529a5e769051b4d1372748f6

Observation e92b42ec-b7e8-40d5-9e4b-ac03b7a92cd9 · inbound

From Hallucination to Structure Snowballing: The Alignment Tax of Constrained Decoding in LLM Reflection cites this paper.

From Hallucination to Structure Snowballing: The Alignment Tax of Constrained Decoding in LLM Reflection ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:30:51.495749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:04:32.931596Z digest=sha256:5a1d7328cc7fae8603b23357b4a396276467b152d4c567e884fa9c36e21b8d19

Observation 91f99d7a-2a47-4957-b787-2caae3e80e74 · inbound

Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction cites this paper.

Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:37:36.780083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:48:39.832057Z digest=sha256:33169a00d27518e9e50b81beec64aeaac94efe32e84708348ed1034052b30f68

Observation cd110505-f34c-4fbf-b9db-7a0b6c0de025 · inbound

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling cites this paper.

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:44:22.077461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T07:37:38.371341Z digest=sha256:06c8f8021bd765c4f1c6404e565f8c452f13cef52184138001b37f8fac8e500a

Observation e118551f-6c45-4c67-9cc5-8e24d7e4d013 · inbound

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning cites this paper.

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 107

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:25:41.607755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T05:33:52.027771Z digest=sha256:c80467ec668aa4fc60944cc5824c7104fa9c206eeb0bddbea99587e208b083b1

Observation cb15932c-4aba-4345-b557-31fa45f113f4 · inbound

More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness cites this paper.

More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T01:00:54.813255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:00:54.813255Z digest=sha256:b2920bb7174b3c7561388719aedfec2d620cecf399d255192be4d7d0020bcc9d