Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:46:18.444164Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2506.12801.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:46:18.444164Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6d60c76f-c66c-46c7-b466-3e13008fdfaa · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Qwen Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e52cbcd-5108-4ad9-a602-db47319b4a2b · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e3d288-5a06-457b-8a66-40f6bb9cb86c · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3a463155-2597-4ef2-964b-74a53a875e98 · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Superhuman ai for multiplayer poker
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a43e3e5e-e61f-4e27-ac90-a90c677fda84 · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Adam: A Method for Stochastic Optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85fc7ca2-1648-45de-af0e-6a7b1c8fa35a · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Suphx: Mastering Mahjong with Deep Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b554b433-b37d-4f9a-9428-7b4829f02ce2 · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents GPT-4 Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 594d8a22-23ea-45a4-93fc-9a61eb5e71ca · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Hello gpt-4o
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8a8aaed6-f233-4b68-a2bb-f811d8be521e · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Mastering the Game of Guandan with Deep Reinforcement Learning and Behavior Regulating
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9d0c5ce8-5151-4d6f-b7e1-8045cc1a05b7 · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Proximal Policy Optimization Algorithms
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79dd1ee9-63a7-4d91-bc3e-329bb2afcfba · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Trust region policy optimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 09d1181f-e15e-4e90-95ec-3fbbca9fc126 · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Mastering the game of go with deep neural networks and tree search
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 786e676c-117e-4f70-a09f-c81640a7d0c6 · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Gemini: A Family of Highly Capable Multimodal Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cab5b04-4471-45d2-b4a0-5688e6113dbf · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Td-gammon, a self-teaching backgammon program, achieves master-level play
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4dcc26b3-738a-4677-b8d4-a528b7371ddd · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2e85127-3aad-4059-b998-c8d48157c9be · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Tree-of-thought prompting for large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4d98cd54-82a1-42b4-90d4-ae650af919ba · outbound
Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents DouZero: Mastering DouDizhu with Self-Play Deep Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.