Pith. sign in

Paper Citation Record · LEDGER

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 5 inbound Pith citation observations for arXiv:2502.06060.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06060 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:57:01.718096Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:00:53.877395Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T07:32:09.425973Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact6
  • verified fuzzy5
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3a65100-2e0f-424c-889b-d3855b56bad2 · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.572507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.572507Z digest=sha256:a27417639ce8ab6543fa48a80410c3b93c9ce1d8b37930dd3d1ccd9dad01b8a8

Observation 82269cb0-6c43-4206-b60e-bb0349a54d42 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 2

Resolution
verified exact
raw_fallback, observed 2026-08-08T16:57:02.161295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.577352Z digest=sha256:95519533e1bcc9f3d06afa8c86dfb3d4b1d377339adea3bc40cb3f6bd9fbcebe

Observation d1411885-52ba-484b-9954-f2d1d1bc9662 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.581162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.581162Z digest=sha256:dfc984d00a603be7f05bcb477e95456036d98358737fc3b4eff1631f089e48e0

Observation 803c69ef-361d-4977-b95c-e6f653a82a6c · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-08T16:57:02.062627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.585402Z digest=sha256:6ac976e48c7bfc39481cb413b5c3d5bd75ae1622561bd6013547fcec56ec171e

Observation bb82cba9-8dc6-4512-afd3-3031fbd2d34a · outbound

This paper cites Ho, Thomas L.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Ho, Thomas L

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:02.447577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.589375Z digest=sha256:0727a5f1f57d019a9639f1754a4dee00d93ba9b850954bfe518dec4f304ad1cf

Observation 94c093dc-24a3-4329-b144-fe9175949061 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.435082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.592853Z digest=sha256:9543b2f498f6b992141c3002235a9d9e390f8992c7724a2cc0dbc5d12e1aeccd

Observation a4a05442-3376-414f-8d4f-72ea7717feac · outbound

This paper cites Optimizing Robotic Manipulation with Decision-RWKV: A Recurrent Sequence Modeling Approach for Lifelong Learning.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Optimizing Robotic Manipulation with Decision-RWKV: A Recurrent Sequence Modeling Approach for Lifelong Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-08T16:57:01.906110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.596395Z digest=sha256:9ff61ac7a597aee78e024f1647de1973c6010d9c1ee5a1779aa48e6ecd3d9a31

Observation e7b51152-717d-48c8-ab6f-73cf92841f4d · outbound

This paper cites Miller, Sasha Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David Wu, Hugh Zhang, and Markus Zijlstra.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Miller, Sasha Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David Wu, Hugh Zhang, and Markus Zijlstra

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.600162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.600162Z digest=sha256:1e1cc89277e72c7939593825406fea68123b9d7ccfcef254e699dbc5ca4d780a

Observation 1ed4ef97-699d-4132-8d05-361ebe8bf570 · outbound

This paper cites Frank and Noah D.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Frank and Noah D

Reference 9

Resolution
verified exact
doi, observed 2026-08-08T16:57:01.770581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.603404Z digest=sha256:4f45c726c092722dcf8972938186ff131549b619b9563b783afeaadc9e032807

Observation c59bec0d-5c7d-4c24-bccc-d42737c7cbe7 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 10

Resolution
malformed identifier
no resolver link, observed 2026-08-08T16:57:01.606357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.606357Z digest=sha256:3a852f605106089272a7704dd9863c4078d5f789f89d8e373ccc383bbe2559ef

Observation 1e5ebcad-1ce4-46a6-8f6a-8bd5b423d22a · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.416100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.609687Z digest=sha256:f10af5cbf56e5f14aee8bac230b5d414c7a916ed01f3fe34738eb7b1402b55f5

Observation a762d21c-7fc4-418c-99c3-c0772f6901bd · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.613230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.613230Z digest=sha256:8220fe428a53796367e0bdc3dc72cdc56b3d778f33c7c7f775c76d53ac59dfea

Observation fc7e3208-0a00-472b-894f-b98b816d5e20 · outbound

This paper cites Other- Play.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Other- Play

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:02.404647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.616461Z digest=sha256:aa4a73749444a33b447019f00f853c2d621436a4df8168fa8f895be853a4b284

Observation 2f7442a0-b8fe-4c37-bea2-a9a0ce36215e · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.393663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.619725Z digest=sha256:b3f7f49a11fb94322e0b29675c17c85824476823fd260649e1159c0f1e8e4d40

Observation d7d2303e-97d9-4ec2-946f-28fed295ec62 · outbound

This paper cites How Well Can a Long Sequence Model Model Long Sequences? Comparing Architechtural Inductive Biases on Long-Context Abilities.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning How Well Can a Long Sequence Model Model Long Sequences? Comparing Architechtural Inductive Biases on Long-Context Abilities

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-08T16:57:01.890742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.623014Z digest=sha256:a0b58166b15dea69eacc816ce205d0ea9a0fa14720b9d4777e9b37430a62d99e

Observation c50d80ce-5814-4147-a54b-62d404ea2b00 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.381313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.626493Z digest=sha256:1aa4e166a99b81aa9e82723dbfe5f7543d7dd73186c0c48258c6fbc89082f039

Observation bdcd4efd-7c2d-42bf-aa95-eef8ab8a9cf4 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.368836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.630031Z digest=sha256:cf0904e5189d8307a2a23231ecd857347565cc26fbe8141aedab6e429dedea98

Observation 8d1c974c-6301-415b-bae0-1b9c816eb016 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.357294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.633097Z digest=sha256:1b80e07a691ace9c91cd9b50278ad2e97e445750888e7992cb9268269d5a0526

Observation d78b4802-27a3-4b01-8284-ce04436d6d68 · outbound

This paper cites Hidden Agenda: a Social Deduction Game with Diverse Learned Equilibria.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Hidden Agenda: a Social Deduction Game with Diverse Learned Equilibria

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T16:57:01.874532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.636494Z digest=sha256:5d183bdcd67e7ef20f9039fac9db446b3b11336b80e77edee61dbc88b38b4e53

Observation b9fab916-f919-4763-9ae0-5c910b4b35ac · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.345422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.641080Z digest=sha256:aaafc2b63ee2906c2debeab7e6ee571449dceb62a336ab32ac158090cc700999

Observation 3635e34f-e469-49e2-8704-34332322c219 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.333846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.643930Z digest=sha256:a8f12710932eb1db037b975e43adb9840d4992c3347bd5c62ea17090078bd43d

Observation f49ff4b5-ac22-446d-8833-4c97c287a483 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.321355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.647151Z digest=sha256:c964f4203f37d25e5d4bbaa13a69554eb6e81f28272a0b2c5e9f7495fc690483

Observation 5f7ae03b-e11f-4eee-93b8-a7beb62a043e · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.308825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.650590Z digest=sha256:05063b68f556df20395eb9ac12583398772e6d292798fdbc1d9fcf8c4851dcef

Observation 2b224264-2c23-4853-b5f8-c2f8c8b868a4 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.653388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.653388Z digest=sha256:4d600a4217fd7ca73f9d9ed4873ba40fa9345a6f97042b35999ad0230746e510

Observation 3995a771-aa8a-4de3-81db-cc6c6b24386f · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.297187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.656478Z digest=sha256:151744c503252cbaca271ed661c32a0646b27846f82f1850d70fe3657131ff9e

Observation a42f230c-1963-4977-9d4d-f9e96780ac98 · outbound

This paper cites Learning to communicate about shared procedural abstractions.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Learning to communicate about shared procedural abstractions

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-08T16:57:01.858399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.659710Z digest=sha256:2ad388d50aa0018d40986910789ece16b8d87d615352fd717823c9f264dbea47

Observation 2661ece3-9c85-4199-9481-737b0e06edeb · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.285251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.663363Z digest=sha256:d251b1ecd7c347f0d564fbe38e2f52764ae2d7192a8ea223817a3fb884712888

Observation 2fa00c6e-339e-46ec-8bd5-6bd3832d717e · outbound

This paper cites Training language models to follow instructions with human feedback.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Training language models to follow instructions with human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.666917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.666917Z digest=sha256:dcf4065af2d042fd428c3cecd7a19f16ef9554eb710f7265e9666ac11c784b2c

Observation 8b0df138-354e-4b88-984c-e67ba61d28ff · outbound

This paper cites O’Brien, Carrie J.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning O’Brien, Carrie J

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:02.273508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.670410Z digest=sha256:6ba1899b7c893effe28e42ad82f638e50fb353563a0f95990ea4208423e03a72

Observation 6edb8572-601d-4e33-a15a-ab35cc3a17dd · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.673484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.673484Z digest=sha256:864206151dbcf138b5f98bdc5282ba4b95b0aa8694fe8af32d734c617fc5078d

Observation 5e912628-25be-47fa-a17c-6dcebc2e7e29 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.254440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.676803Z digest=sha256:1aba0f12ed45979e8eb7fb6f226d568897fcc7c63642e33cecb18d3cee5e18ca

Observation 493521fc-d916-4604-ba17-648c2501faa0 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.680226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.680226Z digest=sha256:d795b92bcd3118521dde4e166d8bdffedb76a3560c1b750ed49032ce891942c4

Observation 2d97236c-f11e-449b-a795-6bae400c00d6 · outbound

This paper cites Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.687133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.687133Z digest=sha256:318828329b8ab0133af4f94d16873799e1c3dc7caba5901edd7aec475c900745

Observation 5d064cd7-dda9-4876-b173-ef880458fcdb · outbound

This paper cites GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.690429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.690429Z digest=sha256:a6a3c8b04bbfc785bf796b6739fea4911c97ff105a98b42de2ccaa7443364020

Observation 19451257-c0e9-4641-92d3-31179d4a4e02 · outbound

This paper cites Attention Is All You Need.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Attention Is All You Need

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.694241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.694241Z digest=sha256:019e5e5175a9577a49e4ccbd3fb81bfea3ea3c7ac5f2bff9db18f27af97ed9a6

Observation a32063cf-a45b-4d49-b935-53e29b1ad566 · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.235261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.697776Z digest=sha256:7261ed5a86c67ece35baa2ee489473081fd208c896998adce832cc4a5bfc1186

Observation a8d5c4cb-100f-4c6b-be19-372bff2584bb · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.223236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.700924Z digest=sha256:c2644f47aa53065c3eec21deb64e6c97aa4ee99a1fc23875644b4e2b3aeffc45

Observation 6a85906f-9f91-4430-a8e7-fd65d033e350 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.707442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.707442Z digest=sha256:9ab053da649ebaf876c75513c22160e9ed32924e1b28296e953e88884db125ae

Observation d3543e8b-a77b-4390-bea8-73c7e09aeff4 · outbound

This paper cites Self-Rewarding Language Models.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Self-Rewarding Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.710953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.710953Z digest=sha256:d4b388063f74bf60fc05897761c5a74335823e217ea2d260f0b673a2af24249c

Observation 7c0458b4-6fb2-4a7d-b06f-00ede8fca473 · outbound

This paper cites wait” in a room until something changes in the environment, or “go.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning wait” in a room until something changes in the environment, or “go

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:02.185947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.718096Z digest=sha256:2d9fa89e1bdbd0de6ddd044f786544cd03b3f5000ed1c7d737ccb9fb6409f337

Observation 7f886fcd-e415-4c2d-8fdb-38584b05a71a · outbound

This paper cites an unresolved cited work.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-08T16:57:02.197996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.714731Z digest=sha256:fa8f77510a83acdea90b9fe1e5efe34063ec6a682b5c5b97ea4f19ed60c1b730

Observation baddf2b3-91c9-4f54-8480-39f1a8382b7b · outbound

This paper cites Proximal Policy Optimization Algorithms.

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T16:57:01.683745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:57:01.683745Z digest=sha256:18e1036fb9d8334840a3fd6689f872d676cc251f3d79a845165db23762b27c4b

Observation d16c14f6-979f-4cee-8499-8609add9293b · outbound

This paper cites In Advances in Neural Information Processing Systems (NeurIPS).

Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning In Advances in Neural Information Processing Systems (NeurIPS)

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:57:02.211287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T16:57:01.704144Z digest=sha256:27f171eaeec0c89b5b4ca308d50f385da4d1194e73d7937f1ac1c9d715ee8f8f

Pith citing papers

Observation 6f4b43c0-30b2-4d6a-aa41-a68e9ba6cd59 · inbound

AI Agent Behavioral Science cites this paper.

AI Agent Behavioral Science Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning

Reference 128

Resolution
unresolved
no resolver link, observed 2026-08-07T11:00:53.877395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:00:53.877395Z digest=sha256:9690efeb8a3baf014a6ff966ae6b581fbe5eccac716075c712c6bf323a783ae4

Observation 6ed57130-317f-43e1-add1-ca4c61bd1a67 · inbound

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models cites this paper.

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:58.121890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:41:58.121890Z digest=sha256:4464b091696bbe73f18b21dcd2361e548898f1153082a71485b8eaa38d9550a8

Observation deb9a075-edd8-439d-a3ea-2abfb0e03a88 · inbound

Bayesian Social Deduction with Graph-Informed Language Models cites this paper.

Bayesian Social Deduction with Graph-Informed Language Models Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:32:09.427999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T07:27:20.027179Z digest=sha256:9e3c232827d08cb6267f77dfac5bc1f0909d18c32f043623697e6063bdeb1eee

Observation fb4baca6-46fd-4e13-ba7d-a435926ffbd3 · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:49.345480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:47:47.051969Z digest=sha256:76d72e0843456cdde624df34110fcd7b2232263f6aba93ce3b77e04c9f4615f8

Observation 5d7dd24d-003d-4756-b280-7027d65ebcb3 · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:40.124769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:40.124769Z digest=sha256:476d3ddb3505caedfe5565bfb86e786d484910e5fcc5218c3f8826285f3685c8