Pith. sign in

Paper Citation Record · LEDGER

Predicting Task Difficulty Without Rollouts

As of 17 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.05797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05797 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:25:20.526555Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0bba2fe-0bb1-491b-834e-ac0d945a7719 · outbound

This paper cites Olmo 3.

Predicting Task Difficulty Without Rollouts Olmo 3

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.458447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.458447Z digest=sha256:e3a9293e69977bc8c5a642d481aa67b7806155f700538637751185d761960886

Observation 3a9ab401-4ba1-4e9c-b875-53a64c9060c8 · outbound

This paper cites Mle-bench: Evaluating machine learning agents on machine learning engineering.

Predicting Task Difficulty Without Rollouts Mle-bench: Evaluating machine learning agents on machine learning engineering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.893612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T23:25:20.403156Z digest=sha256:560b57e73eb2b3cb3a7533acfcd104b04e3aa2914e0485cdac51a4fb9dd2e74d

Observation 78735c52-245e-4f5a-84b6-7a589b4e3369 · outbound

This paper cites Demystifying prompts in language models via perplexity estimation.

Predicting Task Difficulty Without Rollouts Demystifying prompts in language models via perplexity estimation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.868696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T23:25:20.419707Z digest=sha256:3d3cd019a05afe043d89b2365fb17a7a3992b124cc74f28d1ff29f01749688c6

Observation a34605f4-7e0a-4eb8-8c5d-973f3a906b82 · outbound

This paper cites The Llama 3 Herd of Models.

Predicting Task Difficulty Without Rollouts The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.423485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.423485Z digest=sha256:c6555c82434a9608e082290acccd075121b4510193c61d0d63ecce140b4cff31

Observation e95e670e-c795-4c33-a69c-fdad96b649c1 · outbound

This paper cites A rosetta stone for ai benchmarks.arXiv preprint arXiv:2512.00193,.

Predicting Task Difficulty Without Rollouts A rosetta stone for ai benchmarks.arXiv preprint arXiv:2512.00193,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.431316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.431316Z digest=sha256:7407ddbee94afc9d4154c13171f07d72a41a1b3ae781279ccf0932f07f042f65

Observation e6404e38-164e-44a2-ad34-aeff77ee642e · outbound

This paper cites Auto-Encoding Variational Bayes.

Predicting Task Difficulty Without Rollouts Auto-Encoding Variational Bayes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.438878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.438878Z digest=sha256:1fb142060aa9488ea84ed26722381ecc7d06f30cff18cee063b96a39fa6118d4

Observation a9a3e6b1-6af2-40f7-8702-d757e1c18d07 · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Predicting Task Difficulty Without Rollouts Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.442787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.442787Z digest=sha256:69d3965cd6bdd5db209602dd1fc4e999b4868a05be9c013b02328043b17f6d14

Observation 0d40531c-cc6b-4a43-af1f-56815b822072 · outbound

This paper cites BRIDGE: Predicting Human Task Completion Time From Model Performance.

Predicting Task Difficulty Without Rollouts BRIDGE: Predicting Human Task Completion Time From Model Performance

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:21.501974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T23:25:20.446641Z digest=sha256:f18c1f334068a10d3f1c995b26e1983e737f4929c22ca184110c90d2caaa3fe4

Observation 52501f48-9867-4a49-9d26-9ebe4f115ba3 · outbound

This paper cites Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning.

Predicting Task Difficulty Without Rollouts Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.450375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.450375Z digest=sha256:e4e6b4649a4da4643ba0fbb5add8f3b4af5d9aaa933152862191ad0c6512e17f

Observation d0953aa5-4aa1-4722-8080-8571348054dc · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

Predicting Task Difficulty Without Rollouts GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.466023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.466023Z digest=sha256:78b6a4fcecd2fae03c6c1ce3b0e1d8ff295127845f20acc638bc37a6cb2932d6

Observation 96e45a33-36d2-4edf-9b49-4e1945035ed4 · outbound

This paper cites Humanity's Last Exam.

Predicting Task Difficulty Without Rollouts Humanity's Last Exam

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.470081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.470081Z digest=sha256:07d94246e884fa6df9f69646626c90356f98aa47f613e578681869efe082de36

Observation c6b5a2d5-e75f-492d-a705-1f756d3bd4c2 · outbound

This paper cites Automatic Curriculum Learning For Deep RL: A Short Survey.

Predicting Task Difficulty Without Rollouts Automatic Curriculum Learning For Deep RL: A Short Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.474015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.474015Z digest=sha256:3f9ff65d2c3a6eea1317cf8f3655dc5204fe80be00b9dca5958b1f218ffb7c02

Observation 174efcb9-77d6-4e48-86a6-76dfe68948ad · outbound

This paper cites Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators.

Predicting Task Difficulty Without Rollouts Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:20.966619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T23:25:20.485277Z digest=sha256:f7c6090abf9ccbd1c0189cde82090aaf637b4003cbcfc3199c3edbacdd64557b

Observation 16b42bd2-0ff4-4ff3-a43f-fda7207f427b · outbound

This paper cites Reliable and Efficient Amortized Model-based Evaluation.

Predicting Task Difficulty Without Rollouts Reliable and Efficient Amortized Model-based Evaluation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.489045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.489045Z digest=sha256:65142f26a56b5581e5cfe19e389e1b7bc2d8f3b8c8e11ff3ca44e92f920a01cb

Observation a9d53eba-cd6f-4842-b84f-b86fbc306157 · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

Predicting Task Difficulty Without Rollouts Benchmark Data Contamination of Large Language Models: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.498536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.498536Z digest=sha256:84a05424ae46597a97e5f73428958aadc0e04495383852bb71d001059f97e6a9

Observation 0c8d5afb-8c2e-454f-9cf2-273d3799e70d · outbound

This paper cites An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382,.

Predicting Task Difficulty Without Rollouts An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.503466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.503466Z digest=sha256:cb9392a85dda6f91c5cfab6b8a423d42ff704fdfed4034833b7d9aff631f3358

Observation 708b96d4-f535-475c-b762-3cec4bfe223d · outbound

This paper cites Qwen3 Technical Report.

Predicting Task Difficulty Without Rollouts Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.507412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.507412Z digest=sha256:30cecf4685f09ffbc85f1698268938104e6bcd65c4afe568a5e2ef3c7c7ac07e

Observation 311b4df5-1aed-4769-ae55-d894cc6527ca · outbound

This paper cites Cybench: A framework for evaluating cybersecurity capabilities and risks of language models.

Predicting Task Difficulty Without Rollouts Cybench: A framework for evaluating cybersecurity capabilities and risks of language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.851248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T23:25:20.511381Z digest=sha256:a4358bea9cef91781701d994f3e37564af1ab77863aff719400f52729fc69eb8

Observation 05eabc8a-2683-4b64-944c-a0883605f3e2 · outbound

This paper cites Edis: Diagnosing llm reasoning via entropy dynamics.arXiv preprint arXiv:2602.01288,.

Predicting Task Difficulty Without Rollouts Edis: Diagnosing llm reasoning via entropy dynamics.arXiv preprint arXiv:2602.01288,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.515298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.515298Z digest=sha256:89bc02031b385579db7d6b55a0005b4bfd7c8821a171079cf0ea17f1b9d9ce41

Observation 2911df35-45ae-4891-8e95-dc49273de9e2 · outbound

This paper cites an unresolved cited work.

Predicting Task Difficulty Without Rollouts Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:25:21.839752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T23:25:20.518691Z digest=sha256:ba6fb894a134d0f818d878f3008ed0196c7752331fe1be8d6333802fdb159795

Observation 43ab6abf-30d6-46cc-aa29-6f637b18a21b · outbound

This paper cites A.2 IRT fitting details The IRT model is fit on the observed binary entries of the agent-task response matrix.

Predicting Task Difficulty Without Rollouts A.2 IRT fitting details The IRT model is fit on the observed binary entries of the agent-task response matrix

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.828224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T23:25:20.522497Z digest=sha256:0db0d6eaae64f8eea68384533f7381d7bfe5a3c006a77ae1843245c6f7496439

Observation f09895f7-b821-4cbf-aa18-3855e666cc36 · outbound

This paper cites an unresolved cited work.

Predicting Task Difficulty Without Rollouts Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:25:21.815890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T23:25:20.526555Z digest=sha256:b1e929927ce2e8a5c7541d7e96fdd5d1656b3ae23a7cd189a167b4a24003558a

Observation 448f91d0-4001-4474-ba87-47060b70a1fc · outbound

This paper cites HCAST: Human-Calibrated Autonomy Software Tasks.

Predicting Task Difficulty Without Rollouts HCAST: Human-Calibrated Autonomy Software Tasks

Reference 1960

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.478107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.478107Z digest=sha256:eede72d9a65c4ecadaefee3107819b2e62da9332569f4441f263514e73acf0f7

Observation fd24008f-7d94-46d1-b45c-723447228ba8 · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Predicting Task Difficulty Without Rollouts RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.492888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.492888Z digest=sha256:699d4bd86bce87390c871b7e5b662208fd1b22258c7938054cd19acb69004c63

Observation 4287b8ed-1894-4439-9cc8-db77e16e5bd2 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Predicting Task Difficulty Without Rollouts Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.454221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.454221Z digest=sha256:6641c695dff55e9c604c79bd2c2e8562f84dd76c14505d8ee498fc2366038da4

Observation aa97920f-3c40-4c05-9308-fa9bbbfcba3b · outbound

This paper cites Refining Minimax Regret for Unsupervised Environment Design.

Predicting Task Difficulty Without Rollouts Refining Minimax Regret for Unsupervised Environment Design

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.394203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.394203Z digest=sha256:de6cd8023e123ee60052b01fb0a04fb017399578c987a3de633b67c89fe5e2fc

Observation 73b60335-e28b-496b-b9b9-fc4add245555 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

Predicting Task Difficulty Without Rollouts Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.434930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.434930Z digest=sha256:6a9ac5135e01d9ef055bf0bfd62da2089658b9c792884a7b141acbc23d9f4162

Observation fe6ff6e8-8897-44bf-9a39-46a31b45b999 · outbound

This paper cites LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking.

Predicting Task Difficulty Without Rollouts LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.427524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.427524Z digest=sha256:4141da39f268a001830bce45c3cc4f5ffe859c946e66d375cd2bce73f791a95c

Observation 3a2f6e1b-4f0c-4f52-8d32-bb4b8e9c97af · outbound

This paper cites Agent psychometrics: Task-level performance prediction in agentic coding benchmarks.

Predicting Task Difficulty Without Rollouts Agent psychometrics: Task-level performance prediction in agentic coding benchmarks

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.415444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.415444Z digest=sha256:251677987b0bef1247295203971be8a830ad575e75747ff995e23057254db80b

Observation e07f15b2-fd35-44d4-8798-77fedaa9da41 · outbound

This paper cites Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models.

Predicting Task Difficulty Without Rollouts Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.881241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T23:25:20.411400Z digest=sha256:c4f5f48e64c6809a7d1df3bf89085947219768c652c31dc82e728411eff097cc

Observation b96fd2cc-786b-4f1f-806b-5527f438c682 · outbound

This paper cites Soft contamination means benchmarks test shallow general- ization.arXiv preprint arXiv:2602.12413,.

Predicting Task Difficulty Without Rollouts Soft contamination means benchmarks test shallow general- ization.arXiv preprint arXiv:2602.12413,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.481760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.481760Z digest=sha256:08b85348a1f7f56a4839e9d38ebe5d556f11b108b16aaba3786226b80560521f

Observation c1dd55f9-450c-4c5f-9925-693c8860ec52 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

Predicting Task Difficulty Without Rollouts SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.407209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.407209Z digest=sha256:406521b844c78e639c8682379a4ee3d93cc7b95ba8b32ca005349497afb19114

Observation 25939d38-e636-4860-88cd-3a9c09926ee3 · outbound

This paper cites The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?.

Predicting Task Difficulty Without Rollouts The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:21.774648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T23:25:20.398733Z digest=sha256:f3990871d4f3a5cfa43c360207244f1ebc31d2eff82ed6cf80614d945c8a823e

Observation a2a1c9aa-b064-4145-a9b9-eb1f5ed5b890 · outbound

This paper cites Attention head entropy of llms predicts answer correctness.arXiv preprint arXiv:2602.13699,.

Predicting Task Difficulty Without Rollouts Attention head entropy of llms predicts answer correctness.arXiv preprint arXiv:2602.13699,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.462141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.462141Z digest=sha256:f8ae3fa51311fd368170b7cf1d5189437e7e8a2c35de1934e02abf3bb5a58bc7

Observation a0bd2f59-d334-4cf0-b7c9-ef94231b0540 · outbound

This paper cites How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks.

Predicting Task Difficulty Without Rollouts How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.389296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.389296Z digest=sha256:32a115f34cb0c4cc25ae218bf2c31c5fb2cc241f26074924ab5143ce5482a60d

Pith citing papers

No inbound Pith citation observations are available.