Pith. sign in

Paper Citation Record · LEDGER

LongReasonArena: A Long Reasoning Benchmark for Large Language Models

As of 22 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:2508.19363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19363 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:51:32.176410Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T17:12:24.565155Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T17:15:51.605247Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 11ef3de4-451f-432f-94a9-e05fefabb761 · outbound

This paper cites L-eval: Instituting standardized evaluation for long context language models.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models L-eval: Instituting standardized evaluation for long context language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.861933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.008462Z digest=sha256:2966f5fd8139c38a39e3446d329a4be73b927fe9c2a0f11cadd226716aa0f1d0

Observation 83b9d2d6-6fee-483c-b6d4-057613265663 · outbound

This paper cites Claude 3.7 sonnet system card.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Claude 3.7 sonnet system card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.014136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.014136Z digest=sha256:f60fd4fecdf428e9be6ae4088465b4dc7b7c0321b1ac97e700396aa8df24187d

Observation f04d89fb-afaa-4542-8e93-f7255cdc5847 · outbound

This paper cites Program synthesis with large language models, 2021.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Program synthesis with large language models, 2021

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.019484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.019484Z digest=sha256:eae888126ac002cb4f2f49ad0cb832ec52223601a5f570b04106240a6cca5e74

Observation 9e3bdcfd-9c47-4dff-ba7d-6eeeda6934bf · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longbench: A bilingual, multitask benchmark for long context understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.817698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.025654Z digest=sha256:ac87412ca6a95ddb5c4e6278416801af63c26a2ae3ac131e0585078ebbf1238f

Observation 37d218c6-2947-44f0-99b2-e246952933e1 · outbound

This paper cites an unresolved cited work.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.031629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.031629Z digest=sha256:28057a4150da4c4e87f07a713e531cbafdcaed9f01c1d5aef8853dcd2f46d26e

Observation 22942380-c5c6-4ad2-93db-3d973e3a7bd7 · outbound

This paper cites LR B ench: Evaluating long-chain reflective reasoning capabilities of large language models via constraint satisfaction problems.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models LR B ench: Evaluating long-chain reflective reasoning capabilities of large language models via constraint satisfaction problems

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.785049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.036864Z digest=sha256:8b75700cd5f7342e57d239c7ea4d6c849a06ae8f01c3ed25bf59ec9b849d16c9

Observation 9086b855-b637-48ae-8fab-424d8e9deb8a · outbound

This paper cites an unresolved cited work.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.042642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.042642Z digest=sha256:7856119220b65712fd574281cc47389873a94b66927abe69f35cd5822005f000

Observation da4897a4-a170-4645-ba63-efc47f0ed0a0 · outbound

This paper cites Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.755102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.048204Z digest=sha256:d6fa316a420a62c7683d607e19c076cded560e55132256eedcebcb19433c78a8

Observation b56c4e72-9614-4507-9c6a-3c851a8248c7 · outbound

This paper cites The Llama 3 Herd of Models.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.053211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.053211Z digest=sha256:b815392ae0297eb97ff287dba721427af879efdc3a2786e2fdebe787fe48bd78

Observation 9d3662b6-92dc-4e8f-bda7-ff9e42432c56 · outbound

This paper cites Measuring coding challenge competence with apps.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Measuring coding challenge competence with apps

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.738819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.058557Z digest=sha256:c76fa986ccf8da1fc3e9aa9f9b17713ad606efd463130d973f55335fb99d6f61

Observation 007f3dc9-ac80-485e-b61e-b99be5be8f9c · outbound

This paper cites Ruler: What’s the real context size of your long-context language models? In First Conference on Language Modeling , 2024.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Ruler: What’s the real context size of your long-context language models? In First Conference on Language Modeling , 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.721688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.063645Z digest=sha256:ec0f037abe708c09e78a5c63eede8590e4bba68c6c95b6da8d4c5280805b634c

Observation 8e8f6709-e8a3-46eb-8d97-95febb8296d2 · outbound

This paper cites Qwen2.5-Coder Technical Report.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Qwen2.5-Coder Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.069073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.069073Z digest=sha256:bd3d51e557c82e512e51b0b32a6cb5d8cdf637a650107a5d2a5f33bc147246bb

Observation aa31b8e0-2957-483c-b73d-525ff64a29ce · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.705364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.074814Z digest=sha256:4561f71007c70d72c53505ad47ba5351301b4b0b68bdb07b5ec94c8b9166d0b0

Observation 82128b88-5087-4324-a59e-86ccf3fc497a · outbound

This paper cites Needle In A Haystack - pressure testing LLM s.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Needle In A Haystack - pressure testing LLM s

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.687647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.079513Z digest=sha256:932670bc651d1f5b9a111c1dd6ad86c41c61d465ae1de1d19d8f80260c5192fc

Observation cde42c8e-5be1-4a33-ad16-f3fdeee6ab84 · outbound

This paper cites Babilong: Testing the limits of llms with long context reasoning-in-a-haystack.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Babilong: Testing the limits of llms with long context reasoning-in-a-haystack

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.671975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.084762Z digest=sha256:dca28a73bb5f7d0d075544ece02a0e5653efc0ad59e3af25bac453c211511796

Observation 587cae72-04bb-4b6c-a35d-1a70eee7464c · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Efficient memory management for large language model serving with pagedattention

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.655989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.089749Z digest=sha256:e1e58e87d57cde7aab2a3bbc5b3155ab34859efe7b45424657c71dc65ee9d544

Observation 7369c73c-db9a-48ec-82b6-ce5e909688fc · outbound

This paper cites Longgenbench: Long-context generation benchmark.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longgenbench: Long-context generation benchmark

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.639951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.094827Z digest=sha256:a61c29d4fe6580f51bd18e8d9750287c01dc631485e320d5f086e2913ac2eb6e

Observation d530507a-6eb7-49d2-8df3-d70b33244b23 · outbound

This paper cites Codei/o: Condensing reasoning patterns via code input-output prediction, 2025.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Codei/o: Condensing reasoning patterns via code input-output prediction, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.622635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.100330Z digest=sha256:6ee151958fdfef911ac4279587e907af607076c9d7345931892dd0065ceebb3d

Observation afd5edd6-3041-45d2-a19e-bfe4605685c2 · outbound

This paper cites Longreason: A synthetic long-context reasoning benchmark via context expansion.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longreason: A synthetic long-context reasoning benchmark via context expansion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.106332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.106332Z digest=sha256:f5174c7287beb1ac664fdad41ac9aff1f3075d6482e3d6cc70252c8e22cde6ed

Observation fb36d692-747d-4768-a391-949e6895bfd3 · outbound

This paper cites ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.111596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.111596Z digest=sha256:120fba76c5e3a9b8ba2d52f174f2ac8f212c01c76ff97948986f0cc478f84223

Observation 0bbb19dd-f518-456a-895a-dd56813da96e · outbound

This paper cites Large Language Models as Code Executors: An Exploratory Study.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Large Language Models as Code Executors: An Exploratory Study

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.116702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.116702Z digest=sha256:7d7a71c3b69d052a0e89723eeda441319dd288357e19498467a221738de3b150

Observation 62d65c44-5632-4c15-96dd-8b952bc853cf · outbound

This paper cites GPT-4o System Card.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models GPT-4o System Card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.121877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.121877Z digest=sha256:120333aa4fe18ed0834fc2d6f65a079821f064c11470f590f125accd2af93b13

Observation 07dddfad-b035-405d-9d03-1cec5bda189b · outbound

This paper cites OpenAI o1 System Card.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models OpenAI o1 System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.127034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.127034Z digest=sha256:c714769dcdda52a974fdc8cfc118b2dfdc7493080574d242e722d2740999abed

Observation 68c2adb3-59fc-4371-99b8-a3b3b0b3d0e8 · outbound

This paper cites QwQ-32B : Embracing the power of reinforcement learning, March 2025.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models QwQ-32B : Embracing the power of reinforcement learning, March 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.606333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.132367Z digest=sha256:510da1a6ab69d7ab2b9c721a3216727462e977e997093623cd72859ac232c7bb

Observation 51b21b92-873c-4ea0-a5cb-cdc7dd413440 · outbound

This paper cites Longgenbench: Benchmarking long-form generation in long context llms, 2025.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longgenbench: Benchmarking long-form generation in long context llms, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.585678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.137583Z digest=sha256:9a5838fabe645e9460ab3125bb25b29499af8ba9d75074fa218efb4cd44d6f75

Observation ba883100-1b0f-45d7-b249-02cb092fba8e · outbound

This paper cites Thoughts are all over the place: On the underthinking of o1-like llms, 2025.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Thoughts are all over the place: On the underthinking of o1-like llms, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.566596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.143425Z digest=sha256:85416aa7fe58303f596b4a6330991dddbc72f0e7b9d34437090c8ff4ad4bf798

Observation c6f8f972-9d78-458d-8b70-9aa23eaf3561 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.549943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.148617Z digest=sha256:ded882447e3b86fd39ca046c9492ad72fff5d572868341b387f31e4e83d6c010

Observation 2d3677ff-fa16-463b-99da-667f05f9e955 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Tree of thoughts: Deliberate problem solving with large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.532111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.155937Z digest=sha256:a9b0610017f1b756e24bb8a67b2c2409f139fc325bb8d146622ff7393434a003

Observation ed79a072-bcfb-4a41-bc25-554558b72305 · outbound

This paper cites Qwen2.5 Technical Report.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models Qwen2.5 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.161124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.161124Z digest=sha256:2468f68771cddfef5a05ebb0192ba7bbea007171f124af95a751533fc3b1913d

Observation c9be9688-3096-4da2-9b6f-978494eda5b0 · outbound

This paper cites bench: Extending long context evaluation beyond 100k tokens.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models bench: Extending long context evaluation beyond 100k tokens

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:51:32.515796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.166762Z digest=sha256:144e79fd66128107049749782a0f6cf3679ecf9351110e09058666c5825cc84a

Observation c9c1372d-36af-4af5-b030-0f44c37b9f30 · outbound

This paper cites DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:51:32.274409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-05T15:51:32.171536Z digest=sha256:15400bf76a7dfbca8b0b265c6be11eda6796a939468d1fe7ada51bcf2cbdd35b

Observation 2448123f-d8e5-4f46-ada2-a791e302adcd · outbound

This paper cites GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?.

LongReasonArena: A Long Reasoning Benchmark for Large Language Models GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T15:51:32.176410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:51:32.176410Z digest=sha256:101e957ea35d9d001765237eee002f8ada30ee0ce5191ceb049f8b8fa9ce64ca

Pith citing papers

Observation 90c34bf1-150e-42b7-9342-24f0518e0937 · inbound

Efficient Evaluation of LLM Performance with Statistical Guarantees cites this paper.

Efficient Evaluation of LLM Performance with Statistical Guarantees LongReasonArena: A Long Reasoning Benchmark for Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:00:51.735759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T10:58:40.958435Z digest=sha256:a6719c2c428b111beb2f3b339f21e8a6324140786476c65eab6d7a053dfbd1bc

Observation 64e5264d-9d28-4c1e-91f4-950aeb31e9a3 · inbound

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction cites this paper.

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction LongReasonArena: A Long Reasoning Benchmark for Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T17:15:51.607478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T03:58:32.372896Z digest=sha256:fa3136c7fc1939e22ecf2f93e7e3786494ac341e8b2e81e8948a4df8fcc6f525

Observation b887b421-a269-4587-a575-485520251693 · inbound

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction cites this paper.

Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction LongReasonArena: A Long Reasoning Benchmark for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T17:12:24.565155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:12:24.565155Z digest=sha256:2cf18181683629d0d06aa32d9a54c9700ca732f174d0be4fd54b3d5e3f30ade9