Pith. sign in

Paper Citation Record · LEDGER

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents

As of 16 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2506.12801.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12801 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:46:18.444164Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact2
  • verified fuzzy5
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d60c76f-c66c-46c7-b466-3e13008fdfaa · outbound

This paper cites Qwen Technical Report.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:18.389618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:46:18.389618Z digest=sha256:800e6241023649ce840bc307cccfb8903037e27364fbc05f91bdec5d36c021b8

Observation 5e52cbcd-5108-4ad9-a602-db47319b4a2b · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:18.393519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:46:18.393519Z digest=sha256:4766e5daea5a003b22124ec687f3779a7945001ba5831060134ed3293208ac56

Observation d3e3d288-5a06-457b-8a66-40f6bb9cb86c · outbound

This paper cites Superhuman ai for heads-up no-limit poker: Libratus beats top professionals.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Superhuman ai for heads-up no-limit poker: Libratus beats top professionals

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:46:18.663337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T00:46:18.397147Z digest=sha256:f9c51cff02cc4627432aecdb76e606a6947cf10ef28a3ff8038efc7c7eb0a873

Observation 3a463155-2597-4ef2-964b-74a53a875e98 · outbound

This paper cites Superhuman ai for multiplayer poker.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Superhuman ai for multiplayer poker

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:18.400787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:46:18.400787Z digest=sha256:3b887476744662e89435c7e573e36c57211702f3473759b8ff20ff0a9eca67ec

Observation a43e3e5e-e61f-4e27-ac90-a90c677fda84 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Adam: A Method for Stochastic Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:18.404441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:46:18.404441Z digest=sha256:45f4f8ced36d5932f068ed5be0965d0dcfc090f59c5c142091969493ba2bb1ab

Observation 85fc7ca2-1648-45de-af0e-6a7b1c8fa35a · outbound

This paper cites Suphx: Mastering Mahjong with Deep Reinforcement Learning.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Suphx: Mastering Mahjong with Deep Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:18.409319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:46:18.409319Z digest=sha256:87d1859eab3167233fb55ac01b06471e3b89f399dabafe490cac752aeaac8000

Observation b554b433-b37d-4f9a-9428-7b4829f02ce2 · outbound

This paper cites GPT-4 Technical Report.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents GPT-4 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:18.413161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:46:18.413161Z digest=sha256:fa1c90e9a0d992539597a63b40cbfed40de46543d875751bd8868080ef21ae81

Observation 594d8a22-23ea-45a4-93fc-9a61eb5e71ca · outbound

This paper cites Hello gpt-4o.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Hello gpt-4o

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:46:18.640607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T00:46:18.417085Z digest=sha256:185b0622d043d1e66b40a87726f4f12c6b6c09dc9e74099e920c6f9f5a2a6623

Observation 8a8aaed6-f233-4b68-a2bb-f811d8be521e · outbound

This paper cites Mastering the Game of Guandan with Deep Reinforcement Learning and Behavior Regulating.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Mastering the Game of Guandan with Deep Reinforcement Learning and Behavior Regulating

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:46:18.515643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T00:46:18.420109Z digest=sha256:20c1fa37cab5a61947416fc79a667f91eff22888a460363003981669bff84bbc

Observation 9d0c5ce8-5151-4d6f-b7e1-8045cc1a05b7 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Proximal Policy Optimization Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:18.423124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:46:18.423124Z digest=sha256:31887020f74298f6308c4760b0c76a7c313caecb09681907c7e49127d25dcff5

Observation 79dd1ee9-63a7-4d91-bc3e-329bb2afcfba · outbound

This paper cites Trust region policy optimization.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Trust region policy optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:46:18.623923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T00:46:18.426144Z digest=sha256:2d678f182cda0ae83d8ff0acdc18629dffb77c497493e98b30a89a594f034024

Observation 09d1181f-e15e-4e90-95ec-3fbbca9fc126 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Mastering the game of go with deep neural networks and tree search

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:18.429458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:46:18.429458Z digest=sha256:53e32d4435be0aaa7488181a0ddd2883cf46c65f88fc9ab526508c8f49d1e34d

Observation 786e676c-117e-4f70-a09f-c81640a7d0c6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Gemini: A Family of Highly Capable Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:18.432777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:46:18.432777Z digest=sha256:9dfb1c687da138149f90d824cec81bb0ffaa87289f97d24eaab259e16189123e

Observation 2cab5b04-4471-45d2-b4a0-5688e6113dbf · outbound

This paper cites Td-gammon, a self-teaching backgammon program, achieves master-level play.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Td-gammon, a self-teaching backgammon program, achieves master-level play

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:46:18.601342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T00:46:18.435553Z digest=sha256:8520875694d01e21ea8260f176f5347fbee75c9e7ae9ccf3e69f29b839dc8ea8

Observation 4dcc26b3-738a-4677-b8d4-a528b7371ddd · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:18.438225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:46:18.438225Z digest=sha256:0de3b852ee8ab4c2dea0632b8b3f9c64885f8a00a89d3af108620272ff71f7f5

Observation b2e85127-3aad-4059-b998-c8d48157c9be · outbound

This paper cites Tree-of-thought prompting for large language models.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents Tree-of-thought prompting for large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:46:18.577443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T00:46:18.441071Z digest=sha256:f0fb1ac681b60b0fa3bddd7a4cd1798cb91695eb960e4889fb64c1ff14130002

Observation 4d98cd54-82a1-42b4-90d4-ae650af919ba · outbound

This paper cites DouZero: Mastering DouDizhu with Self-Play Deep Reinforcement Learning.

Mastering Da Vinci Code: A Comparative Study of Transformer, LLM, and PPO-based Agents DouZero: Mastering DouDizhu with Self-Play Deep Reinforcement Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:46:18.479386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T00:46:18.444164Z digest=sha256:c7149f71e08a6696866ad0f264fcf0bcbd559acd7ad82c6eba7b6877bf858702

Pith citing papers

No inbound Pith citation observations are available.