Pith. sign in

Paper Citation Record · LEDGER

Predicting Task Difficulty Without Rollouts

As of 10 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.05797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05797 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:25:20.526555Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0bba2fe-0bb1-491b-834e-ac0d945a7719 · outbound

This paper cites Olmo 3.

Predicting Task Difficulty Without Rollouts Olmo 3

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.458447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.458447Z digest=sha256:68cd74f7a2dd3eee2da4b0d0c4f6c54de55efe5f917371d3864317e850dc455b

Observation 3a9ab401-4ba1-4e9c-b875-53a64c9060c8 · outbound

This paper cites Mle-bench: Evaluating machine learning agents on machine learning engineering.

Predicting Task Difficulty Without Rollouts Mle-bench: Evaluating machine learning agents on machine learning engineering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.893612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:25:20.403156Z digest=sha256:c6f15780733310cec2f535cffadb7f09fa21b6d3023069cbc7901afb36dc5c75

Observation 78735c52-245e-4f5a-84b6-7a589b4e3369 · outbound

This paper cites Demystifying prompts in language models via perplexity estimation.

Predicting Task Difficulty Without Rollouts Demystifying prompts in language models via perplexity estimation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.868696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:25:20.419707Z digest=sha256:2e8d5f310f6623a25ee40ec8a4ddc8997019e82fd718996f39fedae8718e946e

Observation a34605f4-7e0a-4eb8-8c5d-973f3a906b82 · outbound

This paper cites The Llama 3 Herd of Models.

Predicting Task Difficulty Without Rollouts The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.423485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.423485Z digest=sha256:45f766ec6151941e578700a904fa60cb11485ef295b0b8b80ab85c0caa116129

Observation e95e670e-c795-4c33-a69c-fdad96b649c1 · outbound

This paper cites A rosetta stone for ai benchmarks.arXiv preprint arXiv:2512.00193,.

Predicting Task Difficulty Without Rollouts A rosetta stone for ai benchmarks.arXiv preprint arXiv:2512.00193,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.431316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.431316Z digest=sha256:0eadd0cf8e6b0209f18ad3a40d8512d73b600051708a9d4fc062fb414201138c

Observation e6404e38-164e-44a2-ad34-aeff77ee642e · outbound

This paper cites Auto-Encoding Variational Bayes.

Predicting Task Difficulty Without Rollouts Auto-Encoding Variational Bayes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.438878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.438878Z digest=sha256:3daf9f851546bdf51b6ccc5617a853ae0030a325fda1d08ed0ca94aa0316902b

Observation a9a3e6b1-6af2-40f7-8702-d757e1c18d07 · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Predicting Task Difficulty Without Rollouts Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.442787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.442787Z digest=sha256:cdf9bcff24207dca55fb886db2250adfc6d7e08c0ac45d8b152539d277789dcd

Observation 0d40531c-cc6b-4a43-af1f-56815b822072 · outbound

This paper cites BRIDGE: Predicting Human Task Completion Time From Model Performance.

Predicting Task Difficulty Without Rollouts BRIDGE: Predicting Human Task Completion Time From Model Performance

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:21.501974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:25:20.446641Z digest=sha256:f5bfc532c01b8e429814dde218324814f8be5600826801254d203fc53ca5c43e

Observation 52501f48-9867-4a49-9d26-9ebe4f115ba3 · outbound

This paper cites Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning.

Predicting Task Difficulty Without Rollouts Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.450375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.450375Z digest=sha256:d69ac98de0208e2c56ecc8bc677fbeac12e60033734941a0d11954ac0183c0e8

Observation d0953aa5-4aa1-4722-8080-8571348054dc · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

Predicting Task Difficulty Without Rollouts GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.466023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.466023Z digest=sha256:dc40e5a751d2535298be1773f874dafbf653bbb3afa4df6775eb0e0b3bd6b1f7

Observation 96e45a33-36d2-4edf-9b49-4e1945035ed4 · outbound

This paper cites Humanity's Last Exam.

Predicting Task Difficulty Without Rollouts Humanity's Last Exam

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.470081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.470081Z digest=sha256:73594102e62f0158cba9d6c7a14c2558377698ecb845392bdbd1ce3e319bd089

Observation c6b5a2d5-e75f-492d-a705-1f756d3bd4c2 · outbound

This paper cites Automatic Curriculum Learning For Deep RL: A Short Survey.

Predicting Task Difficulty Without Rollouts Automatic Curriculum Learning For Deep RL: A Short Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.474015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.474015Z digest=sha256:1164986a9a793847700d0c2be9a5919321e4b1bcbe0f3c3dc9f7807a03eec7ae

Observation 174efcb9-77d6-4e48-86a6-76dfe68948ad · outbound

This paper cites Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators.

Predicting Task Difficulty Without Rollouts Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:20.966619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:25:20.485277Z digest=sha256:0d3b657553bc312482efdb726c1939701413fc2f839146c9af925d4668e9e5ab

Observation 16b42bd2-0ff4-4ff3-a43f-fda7207f427b · outbound

This paper cites Reliable and Efficient Amortized Model-based Evaluation.

Predicting Task Difficulty Without Rollouts Reliable and Efficient Amortized Model-based Evaluation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.489045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.489045Z digest=sha256:60fbbfbc20abd4d7c224108c397e521e3f3be67eea6a36821b6d7f8d0b0242b4

Observation a9d53eba-cd6f-4842-b84f-b86fbc306157 · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

Predicting Task Difficulty Without Rollouts Benchmark Data Contamination of Large Language Models: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.498536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.498536Z digest=sha256:2b0621beb122968c88e0465b4fa8c49707c2358bb89b093854031909b5c50bd8

Observation 0c8d5afb-8c2e-454f-9cf2-273d3799e70d · outbound

This paper cites An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382,.

Predicting Task Difficulty Without Rollouts An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.503466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.503466Z digest=sha256:ef55d343f294caac3c5e0e14ea8848ae51bf1ab4ebdce131313fff732d072fb8

Observation 708b96d4-f535-475c-b762-3cec4bfe223d · outbound

This paper cites Qwen3 Technical Report.

Predicting Task Difficulty Without Rollouts Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.507412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.507412Z digest=sha256:5dce1c82746aae350973ab8aab7baa107db3696a28c9840ffcbc65ce1ab9757e

Observation 311b4df5-1aed-4769-ae55-d894cc6527ca · outbound

This paper cites Cybench: A framework for evaluating cybersecurity capabilities and risks of language models.

Predicting Task Difficulty Without Rollouts Cybench: A framework for evaluating cybersecurity capabilities and risks of language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.851248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:25:20.511381Z digest=sha256:2e8eed717c6a712fe8b77a553eca4aa69e7001d4e1a616e3700200ba7f7e309d

Observation 05eabc8a-2683-4b64-944c-a0883605f3e2 · outbound

This paper cites Edis: Diagnosing llm reasoning via entropy dynamics.arXiv preprint arXiv:2602.01288,.

Predicting Task Difficulty Without Rollouts Edis: Diagnosing llm reasoning via entropy dynamics.arXiv preprint arXiv:2602.01288,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.515298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.515298Z digest=sha256:3b69eb12d38ffb1bc4de499c73665047604bd1b0da5e9208c1c65a41d0b5fbe4

Observation 2911df35-45ae-4891-8e95-dc49273de9e2 · outbound

This paper cites an unresolved cited work.

Predicting Task Difficulty Without Rollouts Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:25:21.839752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:25:20.518691Z digest=sha256:0a4ddd24b9995e3bf9b3e21347ae3ffbe2bd66c012684e7e13caff4603212e94

Observation 43ab6abf-30d6-46cc-aa29-6f637b18a21b · outbound

This paper cites A.2 IRT fitting details The IRT model is fit on the observed binary entries of the agent-task response matrix.

Predicting Task Difficulty Without Rollouts A.2 IRT fitting details The IRT model is fit on the observed binary entries of the agent-task response matrix

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.828224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:25:20.522497Z digest=sha256:5d5600486baebed445a4ab05b9ca1acb00c9ca779a03d5736b051bc8283c4161

Observation f09895f7-b821-4cbf-aa18-3855e666cc36 · outbound

This paper cites an unresolved cited work.

Predicting Task Difficulty Without Rollouts Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:25:21.815890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:25:20.526555Z digest=sha256:1887aa644aae9ea17949f7f57fd4032749fb43cfa7887bcd31385d1fea74fdd4

Observation 448f91d0-4001-4474-ba87-47060b70a1fc · outbound

This paper cites HCAST: Human-Calibrated Autonomy Software Tasks.

Predicting Task Difficulty Without Rollouts HCAST: Human-Calibrated Autonomy Software Tasks

Reference 1960

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.478107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.478107Z digest=sha256:bae293815924daaed2dce6331628e3174a2cffe0d164b9475ebfbe165a5f1b1e

Observation fd24008f-7d94-46d1-b45c-723447228ba8 · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Predicting Task Difficulty Without Rollouts RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.492888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.492888Z digest=sha256:fc21cc759ddd36ae41c22fccad6b400365e2385360f03ea3ce5af58d453cc13b

Observation 4287b8ed-1894-4439-9cc8-db77e16e5bd2 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Predicting Task Difficulty Without Rollouts Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.454221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.454221Z digest=sha256:8871cdc1f475011383fb6dac107b0339f4f49f919ab92e4d6024ea18b9d5b682

Observation aa97920f-3c40-4c05-9308-fa9bbbfcba3b · outbound

This paper cites Refining Minimax Regret for Unsupervised Environment Design.

Predicting Task Difficulty Without Rollouts Refining Minimax Regret for Unsupervised Environment Design

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.394203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.394203Z digest=sha256:9f84f6787588a162c82a815c9fb47ed036c7dab53128d39fd50527f56886cbd8

Observation 73b60335-e28b-496b-b9b9-fc4add245555 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

Predicting Task Difficulty Without Rollouts Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.434930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.434930Z digest=sha256:e7048ac821c77ba3205dd89f6e8ba40415e232cbc83d4597f1476d0d5e6382de

Observation fe6ff6e8-8897-44bf-9a39-46a31b45b999 · outbound

This paper cites LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking.

Predicting Task Difficulty Without Rollouts LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.427524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.427524Z digest=sha256:5803dedb1481cd4cb3f5e907f3b141b3aac78aa3d92d1d352a4578e3fa1c56c2

Observation 3a2f6e1b-4f0c-4f52-8d32-bb4b8e9c97af · outbound

This paper cites Agent psychometrics: Task-level performance prediction in agentic coding benchmarks.

Predicting Task Difficulty Without Rollouts Agent psychometrics: Task-level performance prediction in agentic coding benchmarks

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.415444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.415444Z digest=sha256:8bea91dbb359fcd4ea889caea3e4fcb52a90b34322c809ea039d45d6454513db

Observation e07f15b2-fd35-44d4-8798-77fedaa9da41 · outbound

This paper cites Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models.

Predicting Task Difficulty Without Rollouts Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.881241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:25:20.411400Z digest=sha256:2b3a11b3344aa3f3d215842b678f1721702e4bdba20097cdb51f66d8cdb81b6f

Observation b96fd2cc-786b-4f1f-806b-5527f438c682 · outbound

This paper cites Soft contamination means benchmarks test shallow general- ization.arXiv preprint arXiv:2602.12413,.

Predicting Task Difficulty Without Rollouts Soft contamination means benchmarks test shallow general- ization.arXiv preprint arXiv:2602.12413,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.481760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.481760Z digest=sha256:7a9c099eee2d1cdc7e2e8e142bbc72a7b4dfadcbf50217182457db3338d2ca6a

Observation c1dd55f9-450c-4c5f-9925-693c8860ec52 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

Predicting Task Difficulty Without Rollouts SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.407209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.407209Z digest=sha256:69b28d6fe470195a5704ff6b0cd6d026dd58927ce8ee9c08141ddce51aa10398

Observation 25939d38-e636-4860-88cd-3a9c09926ee3 · outbound

This paper cites The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?.

Predicting Task Difficulty Without Rollouts The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:21.774648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T23:25:20.398733Z digest=sha256:3de83aa40c5eb408a8621ccb8baef61a6d908e5d40c67fb512754ebbea4d7a0d

Observation a2a1c9aa-b064-4145-a9b9-eb1f5ed5b890 · outbound

This paper cites Attention head entropy of llms predicts answer correctness.arXiv preprint arXiv:2602.13699,.

Predicting Task Difficulty Without Rollouts Attention head entropy of llms predicts answer correctness.arXiv preprint arXiv:2602.13699,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.462141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.462141Z digest=sha256:65c33b437013dc7d1fa9d1465565107af01996cde1ec7036c0f6f21fa19fd09c

Observation a0bd2f59-d334-4cf0-b7c9-ef94231b0540 · outbound

This paper cites How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks.

Predicting Task Difficulty Without Rollouts How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.389296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.389296Z digest=sha256:0106a2f1be0a9e4d262c2e5d5a4db92e2724a138c2667eb3728674ff4579dd87

Pith citing papers

No inbound Pith citation observations are available.