Pith. sign in

Paper Citation Record · LEDGER

Exploration and Exploitation Errors Are Measurable for Language Model Agents

As of 5 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2604.13151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.13151 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:00:24.785343Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T06:01:28.311572Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact9
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eddaaf24-f2a2-446d-af0a-49c46fcb1efe · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

Exploration and Exploitation Errors Are Measurable for Language Model Agents gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:21:02.305466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:5df72ee15700897d72ce06e6abe0b07f8892b4456e2c849e9c60ece26eb285bd

Observation 7dfcaba7-93a1-4d78-88e8-fd8e6ef58ddf · outbound

This paper cites Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:02:09.854882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:561faf62af5e56b3c6413056de761787d76d62a2415995f4633d29d52cf0b417

Observation 5427e764-595e-4a3c-8acb-547afa873d03 · outbound

This paper cites Anywherevla: Language-conditioned exploration and mobile manipulation.arXiv preprint arXiv:2509.21006.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Anywherevla: Language-conditioned exploration and mobile manipulation.arXiv preprint arXiv:2509.21006

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:02.275350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:7e325679a73edd8fae46a9262d1490f0b311e3d966a0c825b67bcefe15998d29

Observation f6a71493-9749-4afd-b47a-4bd7ccdef771 · outbound

This paper cites Should You Use Your Large Language Model to Explore or Exploit?.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Should You Use Your Large Language Model to Explore or Exploit?

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-08T02:03:42.262884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:0166246c64a2f7cbd5e8da5d9f5e3a42c9fc55d243054ae8c1ab0b22b1fd458c

Observation 1d10739c-1e8d-4d61-b8cc-aa51580c8936 · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:15:58.398964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:53d398f7d4db643f36cc84f2f8410c84678b0257ca0fa4fb0d407d37c5f9e283

Observation 109a37ae-ff0a-4a39-a9ca-3e532bbdfab4 · outbound

This paper cites Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:21:02.341364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:722d44d2e86e0bcb9a42bce5002b8722a80e34848753f3cafd61961f5d2ad266

Observation b08a4978-6495-429b-bd9a-a600db9a573d · outbound

This paper cites AutoFlow: Automated Workflow Generation for Large Language Model Agents.

Exploration and Exploitation Errors Are Measurable for Language Model Agents AutoFlow: Automated Workflow Generation for Large Language Model Agents

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:02.368384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:dc07991bc9166799b9918fabbf24c37e001c62d05cd777ecf13594ce533d7b4f

Observation 342b5f46-0949-417f-b797-6b2d68b7ad42 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:21:02.382824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:92d9f1bdeffa147336599f0e72c2c6000af46a6e6d32f80652ba50a3adf3a60e

Observation 9f30f0fe-13a5-4269-9fce-7cffe6a444ce · outbound

This paper cites Expanding LLM agent boundaries with strategy-guided exploration.arXiv preprint arXiv:2603.02045.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Expanding LLM agent boundaries with strategy-guided exploration.arXiv preprint arXiv:2603.02045

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:02.296140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:138d8c6025d4c56be84015084445afebbc5e8d71a892384e963e25d68cd545a5

Observation 5bb6699e-ae8a-44c3-8bfe-f21e5c7da09e · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:21:02.316967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:4794428596c742ea98b90a64f03063875d75a066fecae40c2b6b78c4f320c98f

Observation 623c3e4e-444e-419a-8ece-5adcc634f70c · outbound

This paper cites Under review.

Exploration and Exploitation Errors Are Measurable for Language Model Agents Under review

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T09:46:13.716534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:07ea473c19cb7ad2e226ccb31ee38e82cb9b4c9925f9911643c2db95e4ac983d

Observation d0bdd6f9-7ee4-4c65-92a9-ff63fa95865b · outbound

This paper cites action".

Exploration and Exploitation Errors Are Measurable for Language Model Agents action"

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T09:46:13.713882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:00:24.785343Z digest=sha256:2237896c937c2cfdb712cf15847aaf128c7899fe957c06d62079afe9335a596d

Pith citing papers

Observation c56ea90b-128c-46fe-b281-7972fc5aade6 · inbound

Multi-Agent LLMs Fail to Explore Each Other cites this paper.

Multi-Agent LLMs Fail to Explore Each Other Exploration and Exploitation Errors Are Measurable for Language Model Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T06:01:28.311572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:01:28.311572Z digest=sha256:64032e052d15f0dbdf97a0034f06fea109a979a3ec2f4cfa3ae780a0a5278595