Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T09:54:59.746921Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2607.16244.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T09:54:59.746921Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5e0b9589-1f5e-44e7-8682-293f282d0988 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db367e31-9f41-4189-bada-733e1322f6ca · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Schick, J
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4672ff76-e36b-4238-a1e1-0b0495f3ab47 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e56848-1979-4508-9fbf-5a5d0d4bffe5 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f81821-1a5c-45ed-8553-912ea605aed1 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abb30e47-1b4c-44e4-82f1-4a28994bb51c · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Rafailov, A
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 659ec3d1-7201-433a-b62f-2ed733b1d086 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Proximal Policy Optimization Algorithms
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de019db1-0a0f-422d-b629-52f4bc057c23 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Ouyang, J
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8bf6993-3480-4d9c-a108-19cd85ad1cba · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Constitutional AI: Harmlessness from AI Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation accb0374-4555-4dca-9f11-d210c81d3cff · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9748dc5f-43b2-431b-ab4f-bfdb972010e8 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8552f231-b64f-4184-9694-153af1263a8d · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02923fb9-c267-4ac0-a44d-b3571225cc57 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f529a08d-8637-4fc6-be9b-458095dbcc92 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51e8a000-867c-462d-90a2-f9db3743f27e · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5212ec2-4803-4c08-9810-4956e9292443 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22492163-c967-48bf-8ba8-a027a37bf2e3 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Let's Verify Step by Step
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c161a161-aac7-445d-8093-545ac2ec425e · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 593bcd20-98f7-43bd-999b-8e8ed13b62f7 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5349ea33-a3eb-4837-8369-78b6053284fa · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Qwen2.5 Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a489e9c-d9a7-4688-a9e7-61b1b381a3d3 · outbound
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.