Pith. sign in

Paper Citation Record · LEDGER

LongReasonArena: A Long Reasoning Benchmark for Large Language Models

As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:2508.19363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19363 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:51:32.176410Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T17:12:24.565155Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T17:15:51.605247Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 11ef3de4-451f-432f-94a9-e05fefabb761 · outbound

This paper cites L-eval: Instituting standardized evaluation for long context language models.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models L-eval: Instituting standardized evaluation for long context language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.861933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.008462Z digest=sha256:592c4f2689aa447e6433d8095f844a86c83a855e33f622666a3ee1a8271c7482

Observation 83b9d2d6-6fee-483c-b6d4-057613265663 · outbound

This paper cites Claude 3.7 sonnet system card.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Claude 3.7 sonnet system card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.014136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.014136Z digest=sha256:f60fd4fecdf428e9be6ae4088465b4dc7b7c0321b1ac97e700396aa8df24187d

Observation f04d89fb-afaa-4542-8e93-f7255cdc5847 · outbound

This paper cites Program synthesis with large language models, 2021.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Program synthesis with large language models, 2021

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.019484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.019484Z digest=sha256:eae888126ac002cb4f2f49ad0cb832ec52223601a5f570b04106240a6cca5e74

Observation 9e3bdcfd-9c47-4dff-ba7d-6eeeda6934bf · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longbench: A bilingual, multitask benchmark for long context understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.817698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.025654Z digest=sha256:b1c9eed97db648675ce9e9c104773d7a02ef31c96052b3fa96214ea6636692c0

Observation 37d218c6-2947-44f0-99b2-e246952933e1 · outbound

This paper cites an unresolved cited work.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.031629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.031629Z digest=sha256:28057a4150da4c4e87f07a713e531cbafdcaed9f01c1d5aef8853dcd2f46d26e

Observation 22942380-c5c6-4ad2-93db-3d973e3a7bd7 · outbound

This paper cites LR B ench: Evaluating long-chain reflective reasoning capabilities of large language models via constraint satisfaction problems.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models LR B ench: Evaluating long-chain reflective reasoning capabilities of large language models via constraint satisfaction problems

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.785049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.036864Z digest=sha256:c9f0d75fef661589c05febded18aefd0de3d692905e22d826651db83c4e8c3e8

Observation 9086b855-b637-48ae-8fab-424d8e9deb8a · outbound

This paper cites an unresolved cited work.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.042642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.042642Z digest=sha256:7856119220b65712fd574281cc47389873a94b66927abe69f35cd5822005f000

Observation da4897a4-a170-4645-ba63-efc47f0ed0a0 · outbound

This paper cites Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.755102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.048204Z digest=sha256:4483dd6107b78d09d792163c6ac3d7b662a5156b70ac35f4ae328a8058591e25

Observation b56c4e72-9614-4507-9c6a-3c851a8248c7 · outbound

This paper cites The Llama 3 Herd of Models.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.053211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.053211Z digest=sha256:0895f03593eae76a0dcefc8a8425fa152a8ee7fa52c3cb3c374c05d34191a92e

Observation 9d3662b6-92dc-4e8f-bda7-ff9e42432c56 · outbound

This paper cites Measuring coding challenge competence with apps.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Measuring coding challenge competence with apps

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.738819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.058557Z digest=sha256:0de9adacd238b7e41ec3936a0b5b23328409647927702ccc8553f4cc94a96fe8

Observation 007f3dc9-ac80-485e-b61e-b99be5be8f9c · outbound

This paper cites Ruler: What’s the real context size of your long-context language models? In First Conference on Language Modeling , 2024.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Ruler: What’s the real context size of your long-context language models? In First Conference on Language Modeling , 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.721688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.063645Z digest=sha256:819b1be6056dcf40feb8faf5d410ed63d3722a0d603b8c74295d76705ffe3a23

Observation 8e8f6709-e8a3-46eb-8d97-95febb8296d2 · outbound

This paper cites Qwen2.5-Coder Technical Report.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Qwen2.5-Coder Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.069073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.069073Z digest=sha256:bd3d51e557c82e512e51b0b32a6cb5d8cdf637a650107a5d2a5f33bc147246bb

Observation aa31b8e0-2957-483c-b73d-525ff64a29ce · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.705364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.074814Z digest=sha256:a0fce34cd7bdab2e6abd10c48a21c7b8cca25b976b2a69180e5d2fdc73b56956

Observation 82128b88-5087-4324-a59e-86ccf3fc497a · outbound

This paper cites Needle In A Haystack - pressure testing LLM s.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Needle In A Haystack - pressure testing LLM s

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.687647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.079513Z digest=sha256:ed5cd55bc820715b654fe0cc1146bb4b83f9fa2c70f15de048a9915eedc73694

Observation cde42c8e-5be1-4a33-ad16-f3fdeee6ab84 · outbound

This paper cites Babilong: Testing the limits of llms with long context reasoning-in-a-haystack.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Babilong: Testing the limits of llms with long context reasoning-in-a-haystack

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.671975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.084762Z digest=sha256:fbe1819680ff0c02400d51c3f42d7c6627d43729f601796e7d74b69a6f541267

Observation 587cae72-04bb-4b6c-a35d-1a70eee7464c · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Efficient memory management for large language model serving with pagedattention

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.655989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.089749Z digest=sha256:c4bd147563d892959e202d2059f40c56c937033680496eb46fbbe11511637cc4

Observation 7369c73c-db9a-48ec-82b6-ce5e909688fc · outbound

This paper cites Longgenbench: Long-context generation benchmark.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longgenbench: Long-context generation benchmark

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.639951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.094827Z digest=sha256:f543433f57f47ab8947930f9371b2dc8c50e2d0817a159c0ecb92c479b7457a7

Observation d530507a-6eb7-49d2-8df3-d70b33244b23 · outbound

This paper cites Codei/o: Condensing reasoning patterns via code input-output prediction, 2025.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Codei/o: Condensing reasoning patterns via code input-output prediction, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.622635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.100330Z digest=sha256:f0e089dc5d357f28bc5a439b0733aa88dc60fdad810ccc5f8f22f8f11f3ce708

Observation afd5edd6-3041-45d2-a19e-bfe4605685c2 · outbound

This paper cites Longreason: A synthetic long-context reasoning benchmark via context expansion.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longreason: A synthetic long-context reasoning benchmark via context expansion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.106332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.106332Z digest=sha256:f5174c7287beb1ac664fdad41ac9aff1f3075d6482e3d6cc70252c8e22cde6ed

Observation fb36d692-747d-4768-a391-949e6895bfd3 · outbound

This paper cites ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.111596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.111596Z digest=sha256:90c695e8fbe748fc335be3a14a9a10779ffe6d35bd545ed0654920c17c8d4d75

Observation 0bbb19dd-f518-456a-895a-dd56813da96e · outbound

This paper cites Large Language Models as Code Executors: An Exploratory Study.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Large Language Models as Code Executors: An Exploratory Study

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.116702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.116702Z digest=sha256:700a8bab74e13251a7beac1cd0fd8662634e9b14a367b817bb4470b9710809ad

Observation 62d65c44-5632-4c15-96dd-8b952bc853cf · outbound

This paper cites GPT-4o System Card.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.121877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.121877Z digest=sha256:833ef769b6de836e8d5fc78f52bc1ddce9b594aaa9472c34a917fdaf81766c07

Observation 07dddfad-b035-405d-9d03-1cec5bda189b · outbound

This paper cites OpenAI o1 System Card.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models OpenAI o1 System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.127034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.127034Z digest=sha256:134f03be61d3f83ec5299285de52c5a3425b94263762e193e9fd78bf42abd0a6

Observation 68c2adb3-59fc-4371-99b8-a3b3b0b3d0e8 · outbound

This paper cites QwQ-32B : Embracing the power of reinforcement learning, March 2025.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models QwQ-32B : Embracing the power of reinforcement learning, March 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.606333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.132367Z digest=sha256:64603bd04f9f2ce51a2f46c97e8b0ab43d67e91e4d1f142db75b4b153d9bb118

Observation 51b21b92-873c-4ea0-a5cb-cdc7dd413440 · outbound

This paper cites Longgenbench: Benchmarking long-form generation in long context llms, 2025.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longgenbench: Benchmarking long-form generation in long context llms, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.585678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.137583Z digest=sha256:6e0b2ca5a5053a5e7e3f50ad5265909ed3b3e45a1ff73945c4b02679222224ec

Observation ba883100-1b0f-45d7-b249-02cb092fba8e · outbound

This paper cites Thoughts are all over the place: On the underthinking of o1-like llms, 2025.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Thoughts are all over the place: On the underthinking of o1-like llms, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.566596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.143425Z digest=sha256:2bcb814a5c6368425769e038ce1f6bf8e8ba16543d86f0711069ad73fd69e054

Observation c6f8f972-9d78-458d-8b70-9aa23eaf3561 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.549943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.148617Z digest=sha256:f648beb24db547624a3ab10502c34a4bae3f19dc38bc43e92643e32e545787e5

Observation 2d3677ff-fa16-463b-99da-667f05f9e955 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Tree of thoughts: Deliberate problem solving with large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.532111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.155937Z digest=sha256:a0b99aa5786e3aa5c0109ce727c7cc9b980e046b3f2ab12e82249b8d5f7928d6

Observation ed79a072-bcfb-4a41-bc25-554558b72305 · outbound

This paper cites Qwen2.5 Technical Report.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Qwen2.5 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.161124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.161124Z digest=sha256:ab9c21eb73b95e643a408e31e2ae3a4610c984fef7cba3da74609851769508d8

Observation c9be9688-3096-4da2-9b6f-978494eda5b0 · outbound

This paper cites bench: Extending long context evaluation beyond 100k tokens.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models bench: Extending long context evaluation beyond 100k tokens

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.515796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.166762Z digest=sha256:4d69ca8602c36103be38cc6d8fc601b78a54ceb1553ab9b9455d1c5c719f66bf

Observation c9c1372d-36af-4af5-b030-0f44c37b9f30 · outbound

This paper cites DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:51:32.274409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.171536Z digest=sha256:11aa015851dfd2794ca9c71e867d12b5d03b6c67a892c423e2152eabbf5acb92

Observation 2448123f-d8e5-4f46-ada2-a791e302adcd · outbound

This paper cites GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.176410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.176410Z digest=sha256:597ddc6545cdeb3fd2c4dc270ee5e1a813a0e6d7da6a57e01082da9d05cf3b8f

Pith citing papers

Observation 90c34bf1-150e-42b7-9342-24f0518e0937 · inbound

Efficient Evaluation of LLM Performance with Statistical Guarantees cites this paper.

Efficient Evaluation of LLM Performance with Statistical Guarantees LongReasonArena: A Long Reasoning Benchmark for Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:00:51.735759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T10:58:40.958435Z digest=sha256:4232026b6c1f90376b65a5260ff9a8299b1cc40cced3103c6d13db80d91b7e67

Observation 64e5264d-9d28-4c1e-91f4-950aeb31e9a3 · inbound

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction cites this paper.

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction LongReasonArena: A Long Reasoning Benchmark for Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:15:51.607478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T03:58:32.372896Z digest=sha256:9888c8d897ae545579364d6717c946c6ebca86212c47a73bdc06792c6a862d56

Observation b887b421-a269-4587-a575-485520251693 · inbound

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction cites this paper.

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction LongReasonArena: A Long Reasoning Benchmark for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T17:12:24.565155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:12:24.565155Z digest=sha256:571ea1afb0492992d400c699721b0c26af88e53c63123cf73d25c5a05eea2bf0