Pith. sign in

Paper Citation Record · LEDGER

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2605.29782.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.29782 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T04:26:24.074391Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6e7e5f7-7236-48cd-a1f1-03b73300bfd1 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Offline Reinforcement Learning with Implicit Q-Learning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:43:15.445367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:95cab82157e54f95090a681eedb79ee334269930ec7337df7dbebdf065d6b7e8

Observation 29c8779f-fdda-409e-80f1-508e656c1a1f · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:43:14.669708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:7c047c2f70c2eed0a8aaffc9616b87d0760e0586ad8ca86f99adb4916702c4d9

Observation 0a303672-891d-4b0b-9604-177825d56b39 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T08:43:14.666563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:556552ca70ee9ae254c2275f0a8ce02cb67c1335a1f83d5a73ee36785a8bb732

Observation 5f312830-4cf0-4aa6-affd-cbadde38aa03 · outbound

This paper cites Assumption A.1(Answer Parsing).The final answer for both states is parsed from one action a, which corresponds to one token.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Assumption A.1(Answer Parsing).The final answer for both states is parsed from one action a, which corresponds to one token

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:2e7d607ca7c9bfc4baee0b17c23459ea488b87db55af49a5b885bfbdeea04d4a

Observation d212c651-8b0a-4eaa-b0c5-a1c230eb079b · outbound

This paper cites an unresolved cited work.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:3b769a00f31d7538dcedd48a3f93a37acf2c58edc27d218624a2983ee218634d

Observation 4fab8912-a53d-4dbf-8ea5-5f512bcc4d39 · outbound

This paper cites Thus, the corresponding eigenvalue isλ ∥ = dϵ ∥x∥2 2+dϵ.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Thus, the corresponding eigenvalue isλ ∥ = dϵ ∥x∥2 2+dϵ

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:9ada38fd88f6117db5439ad8186aad36a41ed151868c48ffa7c6f52bde6b84d1

Observation f473ed68-7dda-4f91-8346-1d113bd5bf45 · outbound

This paper cites Suppose we continue the generation of s1 and s2 until (T−1) -th state located at the token index ℓ.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Suppose we continue the generation of s1 and s2 until (T−1) -th state located at the token index ℓ

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:65b61f918c48e9a52fceef43d03f998cdb005d0c60bd74da71421f9b942b5631

Observation c78888f1-c705-46c0-b21e-f0723b0e069c · outbound

This paper cites A General Case In previous section, we only prove the correlation between MinDistanceof hidden states and final reward in a simple case, where.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning A General Case In previous section, we only prove the correlation between MinDistanceof hidden states and final reward in a simple case, where

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:afbb19ce6ff8068c43687ad78e0a677117c8c72032af5964ec1e6738d9ecffda

Observation be727993-136a-4e6c-8cbd-2ead0138ad68 · outbound

This paper cites an unresolved cited work.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:3b0da15b43d2ab76ebefeef4da3b1faaa543b297feae1bc1edb2e73ce88594f2

Observation 7956b08b-ea97-42dc-be55-b03b2198d1ae · outbound

This paper cites an unresolved cited work.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:7669b5bb1faffa28097911f08550f1dc36a5c55ef65c150ba19540a02918365e

Observation 5831e3c3-ec11-4cef-a11d-a72dbd323a95 · outbound

This paper cites To expand Theorem A.11 to general case, we need to re- view these three assumptions.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning To expand Theorem A.11 to general case, we need to re- view these three assumptions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:d1483e102c5f2ccc668d3f3bef1a3c3db275a2e628724f7f94a3913b560737b5

Observation cb323b79-50ff-480c-8f3f-c00ae34c329a · outbound

This paper cites Now, the problem becomes, the relation between the final reward of ˆs2 and s2.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Now, the problem becomes, the relation between the final reward of ˆs2 and s2

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:7c552686612c327868d442b3999c27b59d12809dbbc0949c1d4b7ee9368cc47a

Observation d50afa9f-a2d0-42ad-a1e7-4faa3360f6f2 · outbound

This paper cites , η 2}, the dot product score p1,i is replicated ˆri times in the first block of p2, where ˆri are nonnegative integers satisfyingPη2 i=1 ˆri =η 1.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning , η 2}, the dot product score p1,i is replicated ˆri times in the first block of p2, where ˆri are nonnegative integers satisfyingPη2 i=1 ˆri =η 1

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:8492361c5df1d009a2b722b313e3b4b24be23b1745401a1ee9ee7228818d40db

Observation a9977850-c895-4943-8fe6-87ce659b43e2 · outbound

This paper cites , ℓ}, the entry p2,η1+i−η2 in p2 equalsp 1,i +δ i for someδ i ∈R.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning , ℓ}, the entry p2,η1+i−η2 in p2 equalsp 1,i +δ i for someδ i ∈R

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:c6f5dbf3be02d0951226e17e7c560d5b5c6d3c26bc6d2ec625989266b172addc

Observation 9bc1ee54-db30-4a84-bab3-7340a6b767aa · outbound

This paper cites Without loss of generality, we assume s1’s hidden states Xl 1 ∈R η1,d and s2’s hidden states Xl 2 ∈R η2,d, where η1 > η2.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Without loss of generality, we assume s1’s hidden states Xl 1 ∈R η1,d and s2’s hidden states Xl 2 ∈R η2,d, where η1 > η2

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:a064872689d77739cedbe9d019cfc8759fefdf44c905fa609416329c1b5966d9

Observation 8ec788a2-686d-4ecf-8d56-389ead59fee9 · outbound

This paper cites Now, we need to build correlation between ˆXl 2 and Xl.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Now, we need to build correlation between ˆXl 2 and Xl

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:d7ee98e0ec1cdeb95fd6fef859e34126be9e44d2eac4a9d7aefcc6d69189351a

Observation 0df0f422-faef-42d5-9afb-28ed16c62a59 · outbound

This paper cites important.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning important

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:3b59f76a869635d85ddb729895cbf70242e62bd9844a37f73973bbfe2eeabc74

Observation 710fa810-ac94-409c-bb64-0de3069c356f · outbound

This paper cites an unresolved cited work.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T08:37:07.868266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:99688618f2d2e7d1951966f4c3f132353997758be99ce674bafa1c5d8ed1a240

Observation da1feee1-300c-4718-86d9-f081ce785f5b · outbound

This paper cites The result is presented in Table 10.

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning The result is presented in Table 10

Reference 19

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T08:43:15.447991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T08:37:07.868266Z digest=sha256:5ec26c8957be76ca4d6284ce2ed8d74d796f467307beb389ff2d278b419d357d

Pith citing papers

Observation f36e8105-ebcc-453b-9b13-0eabe99137a5 · inbound

APeB: Benchmarking Personalization Ability of Large Language Model Agents cites this paper.

APeB: Benchmarking Personalization Ability of Large Language Model Agents Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T04:26:24.074391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:26:24.074391Z digest=sha256:97d4a3790b07ac3936ea5389040a5d10ef247fd76dacf7270c94f2cb5083e553