Pith. sign in

Paper Citation Record · LEDGER

Stabilizing Transformers for Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:1910.06764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.06764 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:22.815814Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T00:03:51.932172Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7d366faa-f293-4e19-b567-0636d2d4e3a0 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Stabilizing Transformers for Reinforcement Learning

Reference 226

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:36:18.241286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:bf52e472b04c7c11dc70632fcdc3196cf475833cc0b0dff96b077d1209c2a9f5

Observation 1e28d407-380d-484e-9aca-ba9d95119dfe · inbound

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes cites this paper.

Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes Stabilizing Transformers for Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:22.815814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:22.815814Z digest=sha256:992173c3e7ff1adae12cf90822547dff43000ab4277ae9d6186209065f14227f

Observation 93dcfc35-282e-4bef-ade6-6496f8ff10df · inbound

Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation cites this paper.

Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation Stabilizing Transformers for Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:25.771086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:57:25.771086Z digest=sha256:48741fd1070fcd2ddd9b9ae3f914d4a285785e03ffa7037c58befe7e993d71c8

Observation cdeba35a-1efa-4ffe-afcd-fa9a50b79ee1 · inbound

Learning Ordinal Response Policies in Rank-Based Stochastic Prize-Collecting Games cites this paper.

Learning Ordinal Response Policies in Rank-Based Stochastic Prize-Collecting Games Stabilizing Transformers for Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:22.373877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:22.373877Z digest=sha256:3a09dc8d18680f5d3530457e3ffa723d1529775f008a3a66e48f31390c08de91

Observation 8b9d26ea-4635-4fb4-ba2d-cd05c4e34039 · inbound

Anticipatory Reinforcement Learning: From Generative Path-Laws to Distributional Value Functions cites this paper.

Anticipatory Reinforcement Learning: From Generative Path-Laws to Distributional Value Functions Stabilizing Transformers for Reinforcement Learning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:55:50.050619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:28:16.659393Z digest=sha256:89e64a46f7adf4ce52a4f0862d48c40237ae4329c6163d3f9411409119d9cf14

Observation ed5b41cd-213d-4b87-a42a-c20a3c594f09 · inbound

Belief-State RWKV for Reinforcement Learning under Partial Observability cites this paper.

Belief-State RWKV for Reinforcement Learning under Partial Observability Stabilizing Transformers for Reinforcement Learning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:58:19.948012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T21:56:09.642064Z digest=sha256:1ed6b1567e3d2135cdde89726e5844455e264abe4a5e1150f8d9c7b17706e20a

Observation 35652a3f-a239-47e6-b1fa-9e6dc9a139e7 · inbound

Gated Memory Policy: In-Context Memorization and Adaptation cites this paper.

Gated Memory Policy: In-Context Memorization and Adaptation Stabilizing Transformers for Reinforcement Learning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:41:03.387446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T03:16:21.448753Z digest=sha256:1e59f5541cfbd7eb21656933ebbecbff95a88a0c96fb4f154a81fc171a6f2a8f

Observation 7784cf04-b36f-44b1-8b2f-69fbeb97365a · inbound

Graph Transformers and Stabilized Reinforcement Learning for Large-Scale Dynamic Routing Modulation and Spectrum Allocation in Elastic Optical Networks cites this paper.

Graph Transformers and Stabilized Reinforcement Learning for Large-Scale Dynamic Routing Modulation and Spectrum Allocation in Elastic Optical Networks Stabilizing Transformers for Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:43.192778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T18:47:04.091987Z digest=sha256:3647e4870c0d02e10b1a4345670961297f193fea212a0c98d4606c0488fa25ba

Observation 3983e25e-300e-4ca8-84b5-721b3d45cdd0 · inbound

Graph Transformers and Stabilized Reinforcement Learning for Large-Scale Dynamic Routing Modulation and Spectrum Allocation in Elastic Optical Networks cites this paper.

Graph Transformers and Stabilized Reinforcement Learning for Large-Scale Dynamic Routing Modulation and Spectrum Allocation in Elastic Optical Networks Stabilizing Transformers for Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:03:51.939034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T00:02:01.826213Z digest=sha256:51fa3a30e7743508e4f7ace51049f7da753ea99fb0675121938fa1aa19497744