Pith. sign in

Paper Citation Record · LEDGER

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2507.15788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15788 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:26:42.128580Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T17:42:38.122144Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:28.219193Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8f28327-824b-4bef-a666-a4a54e3e7f69 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 1

Resolution
verified exact
doi, observed 2026-08-06T15:26:42.185748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T15:26:39.695665Z digest=sha256:21fe953ec3cf58357314a016bd3dfe0a1424c4d65c2797f22c869b935d0ea5f3

Observation 1266b3c4-f9f9-4b9f-adc5-e638084623e2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:39.772695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:39.772695Z digest=sha256:1ecb2746c286b4d226813e6f5e142040863c73fee1cbafbc31cc253332321e80

Observation 1036b7b4-9e7c-4dc5-911d-4c16ed9ec821 · outbound

This paper cites Understanding Social Reasoning in Language Models with Language Models.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Understanding Social Reasoning in Language Models with Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:39.871666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:39.871666Z digest=sha256:9987afcf7f2393dea9a4d80915c011a0e523f68daab61d02335a0189f17c023e

Observation 50f9f220-21c6-416f-ba66-10950a797d98 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:39.992368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:39.992368Z digest=sha256:3d0c1812124562851af8efb97fb5a4ed7af98f5b69f8ff7e21e0a2a1181000f1

Observation aed2f5a8-7297-4374-ac2f-20f013621a3b · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.154341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.154341Z digest=sha256:2bfbb61f26a97686d65a63ac3e54ea34ea723fcb286fedaefe69e000f531f7c2

Observation 1e1dbb0e-a190-4d4b-b6bb-4a1d5eb046ee · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Evaluating Large Language Models in Theory of Mind Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.278651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.278651Z digest=sha256:62d96c3580c2da9d5c29d774d3fc290cdbf413b88d84423a1ed5676a8e11ff6e

Observation 4633569d-2835-40d6-a3c2-a03082a308e6 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.426349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.426349Z digest=sha256:da24aaf9c1ef878410a41f68cc316d1b2935df64cafc9b5c19d07309c7beb2de

Observation 96a35b91-4428-4388-bdb4-5174da70c734 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.526455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.526455Z digest=sha256:75447ae78bc23bf217b9d7691870488d0f93a74c492099d5418918b136ea520d

Observation 007e2f8e-39e2-475c-8829-2f9866b79efa · outbound

This paper cites Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.632447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.632447Z digest=sha256:cfc2aaf1ef393e41678dbff1ba1a5dc496a285d7a25e34ee0464efc526a9251c

Observation cc647f6b-0f7c-41b7-bbc8-412662dd788c · outbound

This paper cites Training language models to follow instructions with human feedback.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Training language models to follow instructions with human feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.766509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.766509Z digest=sha256:ee248ccb429cfe6cd63fa0127f6919570f650fbf7de21d9446c82f6f32a71f09

Observation 21b2a006-8d70-423c-b673-5d7f54e56a55 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:40.875592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:40.875592Z digest=sha256:f054b8f788cfafd2c247d72a56c18907104a1e171cc09dd72fe11c70c33b1985

Observation 0b57c805-5331-46eb-bbda-e90ebd9232be · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:26:42.305311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T15:26:40.974094Z digest=sha256:5edd627f82ab8f1f86796379f80252d9e828bc3470f6ab8179e1e91569855b8a

Observation e34b3763-d21e-407d-bba4-35096ec82087 · outbound

This paper cites Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:26:42.220425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T15:26:41.077160Z digest=sha256:9773feda0ff85128f308ebe6019b19463b5b7193346abea3e96765b891cd2309

Observation 1f018110-6188-4648-a5d3-ab972d478bf2 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.120715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.120715Z digest=sha256:fbe8f6fb1202e520e98a5b4d37fe6f2662f5c5fa99cb5206250171f7b3a15da2

Observation 3cbe69e1-c50d-471d-b8ab-7ed36264d93d · outbound

This paper cites Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.221542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.221542Z digest=sha256:f66bddc6d452158ea080603be206c6b96f50edd554e3bc0b3c2dff4a00c777c0

Observation e13adfdd-b3a5-4acc-8178-6319d8ade733 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.350355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.350355Z digest=sha256:dae35119c26e0149ed8bc1d62b5b34ffc3196897329e0c7f5df6a4de76824a78

Observation 6a76be41-bfcd-4bbe-906f-37f8621a2200 · outbound

This paper cites Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.514564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.514564Z digest=sha256:7fb794be95f1c62e41fda2898cf4dbe93f8eeda635efd457e25aa7e28393d570

Observation 922a93a1-2cda-4220-8252-7824287ecfff · outbound

This paper cites Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.639155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.639155Z digest=sha256:aba00225a981606498fbec313a8b02634ccd0a8fbc43faec32f3952c56d8031c

Observation 6a0952fd-084b-4284-bb6d-29914603a735 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:26:42.298530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T15:26:41.749750Z digest=sha256:c9c0a3fecccb28cb9fa26928a4d7d4e8c86c491bf7219690a3ef9d1270f5baf5

Observation f158a8bf-d79e-4c01-afe9-508d080014c8 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.829173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.829173Z digest=sha256:20ba0fe4e672e5193f63259c254b4cefe71b6ebd8a1e373141db7ea9eb239e48

Observation 82e049ef-2111-462a-a260-beab818bf18a · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:41.957990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:41.957990Z digest=sha256:fcd2c060ddfe66bce252ff7359912ce0ba299054ca4c2480443da0994ddaae32

Observation c78e948d-5acd-45a5-b72d-439dffeffdb8 · outbound

This paper cites an unresolved cited work.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:26:42.290762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T15:26:42.039503Z digest=sha256:17627e73e3cf4e331ac198676c77da9bd4c9cb883db1e0724ab324591400af35

Observation d6c9a4e1-f7f4-4430-b86c-23dbe1170a8c · outbound

This paper cites online" 'onlinestring :=.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning online" 'onlinestring :=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:42.116273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:42.116273Z digest=sha256:90a178ecf14bbe53f31e5724f49fced25d23a85a35b818cffd3b11c422329093

Observation e6a913be-c6a1-4e59-9562-5535a4dbdd76 · outbound

This paper cites write newline.

Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning write newline

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:42.128580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:26:42.128580Z digest=sha256:b55b99751e3f42d02b9bcfd8fdbce698354569ef13613c250faa60597f97f80d

Pith citing papers

Observation 583dd564-c806-4896-918e-5b6e5914a115 · inbound

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning cites this paper.

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:28.220541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T17:42:38.122144Z digest=sha256:83e5f9768654829e479db0b2af19d34767592bec1ae047139bd35514406ae2af