Pith. sign in

Paper Citation Record · LEDGER

Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2402.06700.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.06700 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:29.215085Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T10:18:11.872878Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f0f69048-4940-4a7c-8c9b-2503ce729b36 · inbound

Advancing Decoding Strategies: Enhancements in Locally Typical Sampling for LLMs cites this paper.

Advancing Decoding Strategies: Enhancements in Locally Typical Sampling for LLMs Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:29.215085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:29.215085Z digest=sha256:c2e9f3c7f0c2bd9280f1382b1bc70253ce378fdcd7558e7b5f6a6243d6a502f7

Observation bcf07401-23ed-448d-a23a-661e2eea50c4 · inbound

Multi-Amateur Contrastive Decoding for Text Generation cites this paper.

Multi-Amateur Contrastive Decoding for Text Generation Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:15.174905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:15.174905Z digest=sha256:7bd9517ebafb8dc503cbf9ef287783e725643930c46086f5da4a7ce99a4def5c

Observation 31b0c102-880e-4755-bbaa-8d84bca1bb8a · inbound

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning cites this paper.

R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T10:01:25.614727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:01:25.614727Z digest=sha256:5ad213f62f6071f695144f4b6f6bd44cf9f4a18ed446298e98ff08fa579ed2e9

Observation 0e5d7203-d1d9-4838-8800-dbbb6a3e972e · inbound

Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation cites this paper.

Pairwise Preference Reward and Group-Based Diversity Enhancement for Superior Open-Ended Generation Entropy-Regularized Token-Level Policy Optimization for Language Agent Reinforcement

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:18:11.875767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T10:15:22.633882Z digest=sha256:5366210161d4da39e3df710cfb5bae69707e4c8a0e2cc27e31e4e37c0feffc73