Pith. sign in

Paper Citation Record · LEDGER

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents

As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2607.16244.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16244 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:54:59.746921Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5e0b9589-1f5e-44e7-8682-293f282d0988 · outbound

This paper cites an unresolved cited work.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.689558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.689558Z digest=sha256:82c0ccdfb78e443f0ec903995b831aca0e32a307311a4c647c02211dc6d623c5

Observation db367e31-9f41-4189-bada-733e1322f6ca · outbound

This paper cites Schick, J.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Schick, J

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.693155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.693155Z digest=sha256:ba64286c9afc56d84275b56e47a9dc1a65e19b6e8e25af85f45d7337c3300cf6

Observation 4672ff76-e36b-4238-a1e1-0b0495f3ab47 · outbound

This paper cites an unresolved cited work.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.696347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.696347Z digest=sha256:e087183464d41cad4257e638a8c067a56ec13bd0e7c0b3938dc7f5fa9ddc18b1

Observation 75e56848-1979-4508-9fbf-5a5d0d4bffe5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.699456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.699456Z digest=sha256:31aea68bc1940f138a7fc96918cbb567f2d2ae01bf99e642c431bf57cf6480f8

Observation 44f81821-1a5c-45ed-8553-912ea605aed1 · outbound

This paper cites an unresolved cited work.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.703091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.703091Z digest=sha256:ccde90b26ca44bfdc74799091735ae1112855561c6e9cebf3f572f3fec48108d

Observation abb30e47-1b4c-44e4-82f1-4a28994bb51c · outbound

This paper cites Rafailov, A.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Rafailov, A

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.706062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.706062Z digest=sha256:af544b5633bfd00a940a35e2bf77c58d6321bead4b54cdb7d4e236f700cce562

Observation 659ec3d1-7201-433a-b62f-2ed733b1d086 · outbound

This paper cites Proximal Policy Optimization Algorithms.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Proximal Policy Optimization Algorithms

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.709186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.709186Z digest=sha256:02b8ee824b0d79ba4bdf7ee80cb56e189335bc6c0dbe274153bcc3de1f7e5a34

Observation de019db1-0a0f-422d-b629-52f4bc057c23 · outbound

This paper cites Ouyang, J.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Ouyang, J

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.712093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.712093Z digest=sha256:be719516448a3e8bc14ac0e71a738f512af7472525264316ee469aa24dbc728f

Observation b8bf6993-3480-4d9c-a108-19cd85ad1cba · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Constitutional AI: Harmlessness from AI Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.714807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.714807Z digest=sha256:8f35257523aea0b7e7e9e760e8c5a2ea72c4fdd1ef5b40f33baab81ca02f5980

Observation accb0374-4555-4dca-9f11-d210c81d3cff · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.717442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.717442Z digest=sha256:e86ef4a57d53370bcffa2e24f35be84020b3069b0508847901500b30dbd35275

Observation 9748dc5f-43b2-431b-ab4f-bfdb972010e8 · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.720373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.720373Z digest=sha256:24cc03b7f82906c4e37f94f93d1db5614715d5a141e1c53f16f9973691c11af8

Observation 8552f231-b64f-4184-9694-153af1263a8d · outbound

This paper cites StreamBench: Towards Benchmarking Continuous Improvement of Language Agents.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents StreamBench: Towards Benchmarking Continuous Improvement of Language Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.723295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.723295Z digest=sha256:d22e778812c605a556311eed590878348278360f55dbedd18cadbea918f15e31

Observation 02923fb9-c267-4ac0-a44d-b3571225cc57 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.726132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.726132Z digest=sha256:e34d32fb1dbd29df0a73ab50d083cafb95216a4ac0198506bc0fb179f23d5933

Observation f529a08d-8637-4fc6-be9b-458095dbcc92 · outbound

This paper cites an unresolved cited work.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.728846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.728846Z digest=sha256:de54a1e2c182076eaf823fb5dbfcdc41de69540f4ece10248abc6ce140c187a2

Observation 51e8a000-867c-462d-90a2-f9db3743f27e · outbound

This paper cites an unresolved cited work.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.731287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.731287Z digest=sha256:095f342724e0d7c17d52a9e2f13b7624d9c0c7dcf237192941ece2b0143baa71

Observation d5212ec2-4803-4c08-9810-4956e9292443 · outbound

This paper cites an unresolved cited work.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.733880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.733880Z digest=sha256:4b2554db1f8f4dd57d62b7d9e68897f0ab0bc57e451d28d8022c80f173ca064b

Observation 22492163-c967-48bf-8ba8-a027a37bf2e3 · outbound

This paper cites Let's Verify Step by Step.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Let's Verify Step by Step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.736381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.736381Z digest=sha256:ae25d29cfcc69752228a9f531633fe1f5146e2a90b855ce39c080d7b0964b4b0

Observation c161a161-aac7-445d-8093-545ac2ec425e · outbound

This paper cites an unresolved cited work.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.739036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.739036Z digest=sha256:fe6ec55083fd1b68102874568667130a47bd6c568547be49b4e21c316225340a

Observation 593bcd20-98f7-43bd-999b-8e8ed13b62f7 · outbound

This paper cites an unresolved cited work.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.741613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.741613Z digest=sha256:0dfd2919a0257133949e5e5792569bdefa84cf672336966e7eef2eb7dd4ad854

Observation 5349ea33-a3eb-4837-8369-78b6053284fa · outbound

This paper cites Qwen2.5 Technical Report.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Qwen2.5 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.743987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.743987Z digest=sha256:76e9a101794a8cf139e2edfeefb680d45bd44ebfaa5db38f618d17145429891a

Observation 9a489e9c-d9a7-4688-a9e7-61b1b381a3d3 · outbound

This paper cites an unresolved cited work.

CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T09:54:59.746921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:54:59.746921Z digest=sha256:9ba0356b05c14e87500c6e937a83e92288df8723e35f210d6621fbc15fd1b332

Pith citing papers

No inbound Pith citation observations are available.