Pith. sign in

Paper Citation Record · LEDGER

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring

As of 13 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2502.07087.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07087 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:53:58.448516Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T08:40:40.910461Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:40:41.541945Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d3d385b-6226-4e31-a870-161a0705d6b8 · outbound

This paper cites write newline.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.343053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.343053Z digest=sha256:46562ddec8a05ec4aa94faab4708a567a61f3011b3c8f4171ee46535abbc1bb5

Observation e1a12ddb-7f55-45ba-ad4f-73451fe84156 · outbound

This paper cites Claude 3.5 Sonnet , 2024.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Claude 3.5 Sonnet , 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.774725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.349218Z digest=sha256:c4c95fd33915ec5c7eea8901b2ef46913c6eb3cdb5a2a4c3df345b787d2399c5

Observation d50af5b8-f711-45f8-98f6-893b9babe3ee · outbound

This paper cites GPT-4 Can't Reason.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring GPT-4 Can't Reason

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.354107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.354107Z digest=sha256:9c25d16f9525a42b5f64d4ca56cb7e9356b82f435c88741ae7a2254992088036

Observation 65fa7dd1-f9fb-4fb8-831e-9e143f8790c4 · outbound

This paper cites DeepSeek-R1 : Incentivizing reasoning capability in LLMs via reinforcement learning, 2025 a.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring DeepSeek-R1 : Incentivizing reasoning capability in LLMs via reinforcement learning, 2025 a

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.761963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.359270Z digest=sha256:72954a301115f910289442bdea4b5f8b078b5d7c11898538f3f8368b8d1d0ab2

Observation 9abc6e86-9f0a-46a1-a9ad-464abcfbd369 · outbound

This paper cites DeepSeek-R1 model card, 2025 b.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring DeepSeek-R1 model card, 2025 b

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.748806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.364348Z digest=sha256:19b5ac772e7077305a256c44b7fdb05cd605feeb7f03cc0421bf9de90d747145

Observation b2b7225e-4314-44fa-8c28-6ac86d1a5b4d · outbound

This paper cites L., Jiang, L., Lin, B.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring L., Jiang, L., Lin, B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.735199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.368996Z digest=sha256:658d4fb26ed54aa1801b8f79c9c3d87ff36615cae9ff7fd1d10038e18621b397

Observation 4326786a-5b26-4ffd-a8d6-14dc41f2c514 · outbound

This paper cites Kimi k1.5 : Scaling reinforcement learning with LLMs , 2025.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Kimi k1.5 : Scaling reinforcement learning with LLMs , 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.721861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.373693Z digest=sha256:e44ea7bca64cbb57d68292b601fe7c12c7d0bbb6e59515bf084dc4d0b75218e9

Observation 6fb7d72f-9a0e-45c9-86f4-ac70a80e85ce · outbound

This paper cites K., Dasgupta, I., Chan, S.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring K., Dasgupta, I., Chan, S

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.708350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.378515Z digest=sha256:3c476a53038b6ab482fa75a7772e8d72d34d15e0cc7285833243e1a59bbf053f

Observation e6082ac6-0cec-4c51-8f52-da3b1e5774ec · outbound

This paper cites Introducing Llama 3.1 : Our most capable models to date, 2024.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Introducing Llama 3.1 : Our most capable models to date, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.694698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.382757Z digest=sha256:65249eef3408f72468bb50c49f95adcc8ee0ca4c2ab50a0487549ea96cd12801

Observation 48f5b795-eb6c-4f2c-90fe-82f8459f3a0c · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.387372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.387372Z digest=sha256:5fb45fd508b52c9f8d3bd8719e55b1906adbb091c8ad83e2a742cb6e7880c28c

Observation 1a570a10-0985-418f-b7c4-58685e3469b8 · outbound

This paper cites A Comprehensive Overview of Large Language Models.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring A Comprehensive Overview of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.392210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.392210Z digest=sha256:1f0af79ed5a7388e5c8a3a050f9e3b5cc4ecdd04784d9b88ac3129a2283dd49b

Observation e0b468e4-69b4-432f-b894-221578b91067 · outbound

This paper cites Hello GPT-4o , 2024 a.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Hello GPT-4o , 2024 a

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.679892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.397984Z digest=sha256:adfcf08531b028e80d95569e53412ef2a0e92017a312c9bdcef76deae8c19469

Observation dd7cb98b-9f09-42b7-852a-0ce6e76eac11 · outbound

This paper cites Learning to reason with LLMs , 2024 b.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Learning to reason with LLMs , 2024 b

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.664477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.402501Z digest=sha256:5b61e608a64f2c4b23dea0609e58aecbab3e4e8a9be11f8a995a87b2fe1bf7f0

Observation 9eace8b9-8428-49f0-a6a7-cacd9842b9de · outbound

This paper cites OpenAI o1-mini , 2024 c.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring OpenAI o1-mini , 2024 c

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.650040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.406833Z digest=sha256:20229484d4573be9d367edd56921495570517af5d329bfd6f85921a5a80888a4

Observation d099ac27-7d05-421c-83e9-4e7be89a6911 · outbound

This paper cites and Hassabis, D.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring and Hassabis, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.636196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.411341Z digest=sha256:33459b6d21f767168f44008530d6a92174f67c05321fc3c8bc21c13194bde885

Observation 434f714e-610f-4443-81eb-c96787744229 · outbound

This paper cites Measuring and narrowing the compositionality gap in language models.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Measuring and narrowing the compositionality gap in language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.622448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.415729Z digest=sha256:9a86927d7ade2a7b5d397477909da0fec70cd2a5383417ba23802b1f84814afc

Observation 308db8f5-f1ff-4f17-9dcf-044cde671c0e · outbound

This paper cites The Butterfly Effect of Altering Prompts: How Small Changes and Jailbreaks Affect Large Language Model Performance.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring The Butterfly Effect of Altering Prompts: How Small Changes and Jailbreaks Affect Large Language Model Performance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.420252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.420252Z digest=sha256:0281627d1b984177dd0ccf4e17046a8e8cc3438f8c2806183cad9c93a45a26cc

Observation 87e53046-c0a7-4f52-81dc-71d90f61234c · outbound

This paper cites H., Sch\" a rli, N., and Zhou, D.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring H., Sch\" a rli, N., and Zhou, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.609040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.425090Z digest=sha256:8b8acb9c6b4e17e33df0de23ff08f4f9fd237f0933233e0b178c65983f2f5946

Observation 7981cc38-c4eb-4601-9ffc-71540aa28923 · outbound

This paper cites On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.429611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.429611Z digest=sha256:43e91074d4d3ea0e1cd8025434ad314b56299b19ddebc90e8a42474a0e29d0d9

Observation a252b1d1-2e84-4f6c-997d-bb3ee9f24002 · outbound

This paper cites LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.434511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.434511Z digest=sha256:24351c26eef120a8d49cb7a1c2019d1377d9873cc532684595e2fa36c173b0c1

Observation 50496835-9d32-4382-ad34-5be903cfd868 · outbound

This paper cites Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.439406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.439406Z digest=sha256:516954d978bf57af5c293a6d5e003a5d408b14811696a19c55505fb16806f7b5

Observation c7cadd4a-96ec-44d4-bab7-8ac4d3c9a951 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Tree of thoughts: Deliberate problem solving with large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.444234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.444234Z digest=sha256:37f20c1b9e4a16d6bc6d340d569884b6b6be9b0e2d64c38b08595e41144697ef

Observation 0f15cfa4-2872-4e42-8026-5218c68e6246 · outbound

This paper cites Larger and more instructable language models become less reliable.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Larger and more instructable language models become less reliable

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.586191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.448516Z digest=sha256:728b6ddd914ccab84dfab3e43d33dfc81e26ea7997f48a0f56451475a4e3c336

Pith citing papers

Observation 89aa5d30-c9ab-4169-ac72-b89355b454db · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring

Reference 259

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.545298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:c256a2a8739cfd905fbe44cf8cbe73fac1f4586c6bdf308a7d8a85016a59bf1b